Resources for Language Technologies
-
EJTN Handbook (Processed)
Handbook on judical training (Processed) This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated...
ZIP (373 amharc) (266 Íoslódálacha)
-
Monolingual corpus from Minutes of the Plenary Sessions of the Croatian Parliament (2016-2018) (Processed)
Minutes of the Plenary Sessions of the Croatian Parliament (2016-2018) were downloaded from http://edoc.sabor.hr . This dataset has been created within the framework of the European...
ZIP (169 amharc) (85 Íoslódálacha)
-
EUIPO - IP case law Italian-English (Processed)
This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action SMART...
ZIP (274 amharc) (175 Íoslódálacha)
-
Bilingual English-Norwegian parallel corpus from Norwegian Maritime Authority website
This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action SMART...
ZIP (453 amharc) (340 Íoslódálacha)
-
Letter of rights for persons arrested on the basis of a European Arrest Warrant (Processed)
Letter of rights for persons arrested on the basis of a European Arrest Warrant (EAW), 1 page, (Processed) This dataset has been created within the framework of the European Language...
ZIP (666 amharc) (557 Íoslódálacha)
-
National Health Fund Dataset (Processed)
The dataset is a 274K-token Polish-English parallel resource in XLIFF format created on the basis of "Diagnosis-Related Groups in Europe" publication of the Polish National Health Fund....
ZIP (345 amharc) (231 Íoslódálacha)
-
Monolingual Greek corpus in the public administration domain
Monolingual Greek corpus, containing 14261776 tokens and 840314 lexical types in the public administration domain. This dataset has been created within the framework of the European...
ZIP (395 amharc) (280 Íoslódálacha)
-
Letter of rights for persons arrested and or detained
Police form, 12 pages. This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation...
ZIP (413 amharc) (296 Íoslódálacha)
-
DA-EN Danish Ministry of Higher Education and Science 2
Parallel texts Danish-English from the Danish Ministry of Higher Education and Science, size 115,000 words, topic: research policy This dataset has been created within the framework of...
ZIP (333 amharc) (209 Íoslódálacha)
-
Portuguese legislation in FR
Portuguese legislation in French (the Parliament's official translations) This dataset has been created within the framework of the European Language Resource Coordination (ELRC)...
ZIP (492 amharc) (372 Íoslódálacha)
-
The Coimisineir Teanga Bilingual Corpus of Reference Documents
General Reference content from the Language Commissioner's Office Size: 6 bilingual Word documents and 44 parallel Word documents This dataset has been created within the framework...
ZIP (339 amharc) (248 Íoslódálacha)
-
The Gaois bilingual corpus of English-Irish legislation
Bilingual corpus of English-Irish legislation provided by the Department of Justice, in two parallel .txt files. Contains 98,758 parallel sentences. This dataset has been created within...
ZIP (432 amharc) (317 Íoslódálacha)
-
Corpus of State-related content from the Latvian Web (Processed)
Latvian Web, home pages of ministries and state public services, army, etc. were crawled, and parallel Latvian-English content was collected. (Processed) This dataset has been created...
ZIP (451 amharc) (346 Íoslódálacha)
-
English-Slovak parallel corpus of texts from The Ministry of Culture of the Slovak Republic
Dataset of various English-Slovak legal texts within agenda of the Ministry, plain text format alligned at the sentence level, the size: 105791 words This dataset has been created within...
ZIP (357 amharc) (249 Íoslódálacha)
-
Convention on the transfer of sentenced persons (English - Greek) (Processed)
Convention, additional protocol on the convention, recomendation R (84) 11 of the Council of Europe, templates on the approval/rejection of transfer requests regarding the convention on...
ZIP (498 amharc) (383 Íoslódálacha)
-
Romanian – English literature corpus
Bilingual Romanian - English literature corpus built from a small set of freely available literature books (drama, sci-fi, etc.). The texts are positionally aligned, i.e. the sentence on...
ZIP (411 amharc) (321 Íoslódálacha)
-
Translation memories from The Ministry of Foreign Affairs of Norway
Translation memories containing translations of EU legislative acts from English to Norwegian Bokmål.
XML PDF ZIP (663 amharc) (540 Íoslódálacha)
-
Corpus RIZIV
Corpus with Dutch and French of the national institute for illness and invalidity insurance
ZIP (625 amharc) (545 Íoslódálacha)
-
English-Estonian corpus from Finnish Information Bank (Processed)
http://www.infopankki.fi - Finland in your language - Information about Finland - Moving to Finland - Living in Finland This dataset has been created within the framework of the European...
ZIP (288 amharc) (186 Íoslódálacha)
-
English-Swedish corpus from Finnish Information Bank (Processed)
http://www.infopankki.fi - Finland in your language - Information about Finland - Moving to Finland - Living in Finland This dataset has been created within the framework of the European...
ZIP (432 amharc) (327 Íoslódálacha)