Resources for Language Technologies
-
Cuimhne Aistriúcháin Ard-Stiúrthóireacht an Aistriúcháin (DGT-TM)
Cuimhne aistriúcháin is ea DGT-TM (abairtí agus na haistriúcháin a cuireadh orthu) atá ar fáil i 24 theanga. Sa chuimhne seo tá píosaí ón Acquis Communautaire, corpas reachtaíochta an...
PDF ZIP (45005 amharc) (4502 Íoslódálacha)
-
COVID-19 multilingual terminology in IATE
The dataset is a collection of multilingual entries related to the SARS-CoV-2 virus and the COVID-19 pandemic, available in IATE, the European Union terminology database. It is a...
Excel XLSX (1490 amharc) (122 Íoslódálacha)
-
IATE
IATE (= “Inter-Active Terminology for Europe”) is the EU's inter-institutional terminology database. IATE has been used by the language services of the EU institutions and agencies since...
HTML JavaScript ZIP (6456 amharc) (6082 Íoslódálacha)
-
Hallituskausi 2007-2011 fi-en
The "Hallituskausi 2007–2011" translation memory is intended for those translating administrative texts between Finnish and English. It includes key policy reports published by the...
XML PDF ZIP (414 amharc) (317 Íoslódálacha)
-
Monolingual corpus from Minutes of the Plenary Sessions of the Croatian Parliament (2016-2018) (Processed)
Minutes of the Plenary Sessions of the Croatian Parliament (2016-2018) were downloaded from http://edoc.sabor.hr . This dataset has been created within the framework of the European...
ZIP (169 amharc) (85 Íoslódálacha)
-
EUIPO - IP case law Italian-English (Processed)
This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action SMART...
ZIP (274 amharc) (175 Íoslódálacha)
-
Monolingual Greek corpus in the public administration domain
Monolingual Greek corpus, containing 14261776 tokens and 840314 lexical types in the public administration domain. This dataset has been created within the framework of the European...
ZIP (395 amharc) (280 Íoslódálacha)
-
Portuguese legislation in FR
Portuguese legislation in French (the Parliament's official translations) This dataset has been created within the framework of the European Language Resource Coordination (ELRC)...
ZIP (492 amharc) (372 Íoslódálacha)
-
The Coimisineir Teanga Bilingual Corpus of Reference Documents
General Reference content from the Language Commissioner's Office Size: 6 bilingual Word documents and 44 parallel Word documents This dataset has been created within the framework...
ZIP (339 amharc) (248 Íoslódálacha)
-
The Gaois bilingual corpus of English-Irish legislation
Bilingual corpus of English-Irish legislation provided by the Department of Justice, in two parallel .txt files. Contains 98,758 parallel sentences. This dataset has been created within...
ZIP (432 amharc) (317 Íoslódálacha)
-
Corpus of State-related content from the Latvian Web (Processed)
Latvian Web, home pages of ministries and state public services, army, etc. were crawled, and parallel Latvian-English content was collected. (Processed) This dataset has been created...
ZIP (451 amharc) (346 Íoslódálacha)
-
Translation memories from The Ministry of Foreign Affairs of Norway
Translation memories containing translations of EU legislative acts from English to Norwegian Bokmål.
XML PDF ZIP (663 amharc) (540 Íoslódálacha)
-
Corpus RIZIV
Corpus with Dutch and French of the national institute for illness and invalidity insurance
ZIP (625 amharc) (545 Íoslódálacha)
-
English-Swedish parallel corpus from the www.visitestonia.com web site
Parallel English-Swedish corpus compiled from the www.visitestonia.com web site by crawling the contents and aligning the parallel data. This dataset has been created within the...
ZIP (317 amharc) (201 Íoslódálacha)
-
Terminology in the domain of Information and Communication Technology (ICT)
Terminology in the domain of Information and Communication Technology (ICT) by Terminology Commission of the Academy of Sciences of Latvia (LAS-TC) This dataset has been created within...
ZIP (267 amharc) (181 Íoslódálacha)
-
Terminology_of_international_contracts_Portuguese_(Processed)
The Portuguese terms extracted from the multilingual terminology of international contracts as provided by the German Foreign Office. Transformed into TBX. 728 terms corresponding to 671...
ZIP (241 amharc) (139 Íoslódálacha)
-
English-Croatian translation memory from the Ministry of Regional Development and EU Funds (Processed)
A translation memory in tmx format with source texts from the Ministry of Regional Development and EU Funds and translations in Croatian by Ciklopea d.o.o. This dataset has been created...
ZIP (323 amharc) (211 Íoslódálacha)
-
Irish Monolingual Corpus from contents of health.gov.ie web site
Irish Monolingual Corpus from contents of health.gov.ie web site This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting...
ZIP (252 amharc) (163 Íoslódálacha)
-
Thematic Vocabulary of Geography (processed)
Thematic Vocabulary of Geography This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated...
ZIP (214 amharc) (131 Íoslódálacha)
-
Citizens Information Bilingual Web-Corpus (Processed)
A web corpus crawled from http://www.citizensinformation.ie. Contains 10,297 parallel sentences of English/Irish that have undergone manual cleaning. May be reproduced and/or re-used free...
ZIP (243 amharc) (156 Íoslódálacha)