TY - GEN
T1 - A workflow for supplementing a Latvian-English dictionary with data from parallel corpora and a reversed English- Latvian dictionary
AU - Deksne, Daiga
AU - Veisbergs, Andrejs
N1 - Publisher Copyright:
© Lexicography in Global Contexts.
PY - 2018
Y1 - 2018
N2 - The lexicon of contemporary languages is changing rapidly, mostly by acquiring new loans and derivations. The change in lexicon is best reflected in the corpora of contemporary languages. Nowadays many collections of parallel-aligned texts are available electronically. To satisfy user needs for a modern, complete, up-to-date dictionary, we created a workflow for enriching the existing Latvian-English dictionary with data from parallel corpora containing lexis commonly used in contemporary language, as well as data from the reversed English-Latvian dictionary. While revising the existing Latvian-English dictionary, we identified some issues, for example, missing feminine forms of the nouns naming nationalities and occupations, representation of the words with optional parts or spelling variations. The task of dictionary improvement was done semi-automatically by the joint work of a lexicographer, computational linguists and programmers. Such natural language processing tools as a tokenizer, part-of-speech tagger, lemmatizer and spell-checker were used to reduce the manual work. As a result, the number of entries has increased by 32%, and the number of translations by 28%.
AB - The lexicon of contemporary languages is changing rapidly, mostly by acquiring new loans and derivations. The change in lexicon is best reflected in the corpora of contemporary languages. Nowadays many collections of parallel-aligned texts are available electronically. To satisfy user needs for a modern, complete, up-to-date dictionary, we created a workflow for enriching the existing Latvian-English dictionary with data from parallel corpora containing lexis commonly used in contemporary language, as well as data from the reversed English-Latvian dictionary. While revising the existing Latvian-English dictionary, we identified some issues, for example, missing feminine forms of the nouns naming nationalities and occupations, representation of the words with optional parts or spelling variations. The task of dictionary improvement was done semi-automatically by the joint work of a lexicographer, computational linguists and programmers. Such natural language processing tools as a tokenizer, part-of-speech tagger, lemmatizer and spell-checker were used to reduce the manual work. As a result, the number of entries has increased by 32%, and the number of translations by 28%.
KW - Electronic dictionaries
KW - NLP tools
KW - Parallel corpora
KW - XML format
UR - https://www.scopus.com/pages/publications/85059433182
M3 - Conference paper
AN - SCOPUS:85059433182
SN - 9789610600978
T3 - EURALEX Proceedings
SP - 127
EP - 135
BT - 18th Euralex International Congress, 2018
A2 - Gorjanc, Vojko
A2 - Krek, Simon
A2 - Cibej, Jaka
A2 - Kosem, Iztok
PB - European Association for Lexicography
CY - Ljubljana
T2 - 18th Euralex International Congress, 2018
Y2 - 17 July 2018 through 21 July 2018
ER -