Resources for Language Technologies
-
Překladová paměť DGT
DGT-TM je překladová paměť (soubor obsahující věty v původním jazyce a jejich překlady od profesionálních překladatelů) ve 24 jazycích. Obsahuje texty acquis communautaire, tj. evropské...
PDF ZIP (45005 zobrazení) (4502 Počet stažení)
-
COVID-19 multilingual terminology in IATE
The dataset is a collection of multilingual entries related to the SARS-CoV-2 virus and the COVID-19 pandemic, available in IATE, the European Union terminology database. It is a...
Excel XLSX (1490 zobrazení) (122 Počet stažení)
-
Romanian – English parallel wordlists
English and Romanian lemmatized wordlists extracted from various resources (including RO-EN Wordnets, the Romanian – English news corpus, the Romanian – English literature corpus, and...
ZIP (885 zobrazení) (765 Počet stažení)
-
IATE
IATE (= “Inter-Active Terminology for Europe”) is the EU's inter-institutional terminology database. IATE has been used by the language services of the EU institutions and agencies since...
HTML JavaScript ZIP (6456 zobrazení) (6082 Počet stažení)
-
EJTN Handbook (Processed)
Handbook on judical training (Processed) This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated...
ZIP (373 zobrazení) (266 Počet stažení)
-
Hallituskausi 2007-2011 fi-en
The "Hallituskausi 2007–2011" translation memory is intended for those translating administrative texts between Finnish and English. It includes key policy reports published by the...
XML PDF ZIP (414 zobrazení) (317 Počet stažení)
-
Monolingual corpus from Minutes of the Plenary Sessions of the Croatian Parliament (2016-2018) (Processed)
Minutes of the Plenary Sessions of the Croatian Parliament (2016-2018) were downloaded from http://edoc.sabor.hr . This dataset has been created within the framework of the European...
ZIP (169 zobrazení) (85 Počet stažení)
-
EUIPO - IP case law Italian-English (Processed)
This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action SMART...
ZIP (274 zobrazení) (175 Počet stažení)
-
Bilingual English-Norwegian parallel corpus from Norwegian Maritime Authority website
This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action SMART...
ZIP (453 zobrazení) (340 Počet stažení)
-
Letter of rights for persons arrested on the basis of a European Arrest Warrant (Processed)
Letter of rights for persons arrested on the basis of a European Arrest Warrant (EAW), 1 page, (Processed) This dataset has been created within the framework of the European Language...
ZIP (666 zobrazení) (557 Počet stažení)
-
National Health Fund Dataset (Processed)
The dataset is a 274K-token Polish-English parallel resource in XLIFF format created on the basis of "Diagnosis-Related Groups in Europe" publication of the Polish National Health Fund....
ZIP (345 zobrazení) (231 Počet stažení)
-
Monolingual Greek corpus in the public administration domain
Monolingual Greek corpus, containing 14261776 tokens and 840314 lexical types in the public administration domain. This dataset has been created within the framework of the European...
ZIP (395 zobrazení) (280 Počet stažení)
-
Letter of rights for persons arrested and or detained
Police form, 12 pages. This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation...
ZIP (413 zobrazení) (296 Počet stažení)
-
DA-EN Danish Ministry of Higher Education and Science 2
Parallel texts Danish-English from the Danish Ministry of Higher Education and Science, size 115,000 words, topic: research policy This dataset has been created within the framework of...
ZIP (333 zobrazení) (209 Počet stažení)
-
Portuguese legislation in FR
Portuguese legislation in French (the Parliament's official translations) This dataset has been created within the framework of the European Language Resource Coordination (ELRC)...
ZIP (492 zobrazení) (372 Počet stažení)
-
The Coimisineir Teanga Bilingual Corpus of Reference Documents
General Reference content from the Language Commissioner's Office Size: 6 bilingual Word documents and 44 parallel Word documents This dataset has been created within the framework...
ZIP (339 zobrazení) (248 Počet stažení)
-
The Gaois bilingual corpus of English-Irish legislation
Bilingual corpus of English-Irish legislation provided by the Department of Justice, in two parallel .txt files. Contains 98,758 parallel sentences. This dataset has been created within...
ZIP (432 zobrazení) (317 Počet stažení)
-
Corpus of State-related content from the Latvian Web (Processed)
Latvian Web, home pages of ministries and state public services, army, etc. were crawled, and parallel Latvian-English content was collected. (Processed) This dataset has been created...
ZIP (451 zobrazení) (346 Počet stažení)
-
English-Slovak parallel corpus of texts from The Ministry of Culture of the Slovak Republic
Dataset of various English-Slovak legal texts within agenda of the Ministry, plain text format alligned at the sentence level, the size: 105791 words This dataset has been created within...
ZIP (357 zobrazení) (249 Počet stažení)
-
Convention on the transfer of sentenced persons (English - Greek) (Processed)
Convention, additional protocol on the convention, recomendation R (84) 11 of the Council of Europe, templates on the approval/rejection of transfer requests regarding the convention on...
ZIP (498 zobrazení) (383 Počet stažení)