Skip to main navigation Skip to search Skip to main content

A workflow for supplementing a Latvian-English dictionary with data from parallel corpora and a reversed English- Latvian dictionary

  • University of Latvia

Research output: Chapter in Book/Report/Conference proceedingConference paperResearchpeer-review

2 Citations (Scopus)

Abstract

The lexicon of contemporary languages is changing rapidly, mostly by acquiring new loans and derivations. The change in lexicon is best reflected in the corpora of contemporary languages. Nowadays many collections of parallel-aligned texts are available electronically. To satisfy user needs for a modern, complete, up-to-date dictionary, we created a workflow for enriching the existing Latvian-English dictionary with data from parallel corpora containing lexis commonly used in contemporary language, as well as data from the reversed English-Latvian dictionary. While revising the existing Latvian-English dictionary, we identified some issues, for example, missing feminine forms of the nouns naming nationalities and occupations, representation of the words with optional parts or spelling variations. The task of dictionary improvement was done semi-automatically by the joint work of a lexicographer, computational linguists and programmers. Such natural language processing tools as a tokenizer, part-of-speech tagger, lemmatizer and spell-checker were used to reduce the manual work. As a result, the number of entries has increased by 32%, and the number of translations by 28%.

Original languageEnglish
Title of host publication18th Euralex International Congress, 2018
EditorsVojko Gorjanc, Simon Krek, Jaka Cibej, Iztok Kosem
Place of PublicationLjubljana
PublisherEuropean Association for Lexicography
Pages127-135
Number of pages9
ISBN (Electronic)9789610600961
ISBN (Print)9789610600978
Publication statusPublished - 2018
Event18th Euralex International Congress, 2018 - Ljubljana, Slovenia
Duration: 17 Jul 201821 Jul 2018

Publication series

NameEURALEX Proceedings
ISSN (Electronic)2521-7100

Conference

Conference18th Euralex International Congress, 2018
Country/TerritorySlovenia
CityLjubljana
Period17/07/1821/07/18

OECD Field of Science

  • 6.2 Languages and Literature

Keywords

  • Electronic dictionaries
  • NLP tools
  • Parallel corpora
  • XML format

Fingerprint

Dive into the research topics of 'A workflow for supplementing a Latvian-English dictionary with data from parallel corpora and a reversed English- Latvian dictionary'. Together they form a unique fingerprint.

Cite this