Skip to main navigation Skip to search Skip to main content

Classifying Multi-Word Expressions in the Latvian Monolingual Electronic Dictionary Tēzaurs.lv

  • University of Latvia

Research output: Chapter in Book/Report/Conference proceedingConference paperResearchpeer-review

Abstract

The electronic dictionary Tēzaurs.lv contains more than 400,000 entries from which 73,000 entries are multi-word expressions (MWEs). Over the past two years, there has been an ongoing division of these MWEs into subgroups (proper names, multi-word terms, taxa, phraseological units, collocations). The article describes the classification of MWEs, focusing on phraseological units (approximately 7,250 entries), as well as on borderline cases of phraseological unit types (phrasemes and idioms) and different MWE groups in general. The division of phraseological units depends on semantic divisibility and figurativeness. In a phrase-me, at least one of the constituents retains its literal sense, whereas the meaning of an idiom is not dependent on the literal sense of any of its constituents. As a result, 65919 entries of MWE have been manually classified, and now this information of MWE type is available for the users of the electronic dictionary Tēzaurs.lv.

Original languageEnglish
Title of host publicationProceedings of the International Conference Computational Linguistics in Bulgaria
Place of PublicationSofia
PublisherInstitute for Bulgarian Language
Pages113-118
Number of pages6
Publication statusPublished - 2024

Publication series

NameProceedings of the International Conference Computational Linguistics in Bulgaria
PublisherInstitute for Bulgarian Language
ISSN (Print)2367-5578

Keywords

  • multi-word expression
  • phraseme
  • phraseological unit
  • idiom
  • semantics

OECD Field of Science

  • 6.2 Languages and Literature

Fingerprint

Dive into the research topics of 'Classifying Multi-Word Expressions in the Latvian Monolingual Electronic Dictionary Tēzaurs.lv'. Together they form a unique fingerprint.

Cite this