Pāriet uz galveno navigāciju Pāriet uz meklēšanu Pāriet uz galveno saturu

Latvian National Corpora Collection - Korpuss.lv

  • Baiba Sau?te*
  • , Roberts Darģis*
  • , Normunds Grūzītis*
  • , Ilze Auziņa*
  • , Krist?ne Lev?ne-Petrova*
  • , Lauma Pretkalniņa*
  • , Laura Rituma*
  • , Pēteris Paikens*
  • , Artūrs Znotiņš*
  • , Laine Strankale*
  • , Krist?ne Pokratniece*
  • , Ilm?rs Poik?ns*
  • , Guntis Bārzdiņš*
  • , Inguna Skadiņa*
  • , Anda Bakl?ne
  • , Valdis Saulespur?ns
  • , J?nis Ziedi?š
  • *Šī darba korespondējošais autors
  • Institute of Mathematics and Computer Science
  • University of Latvia
  • Faculty of Computing
  • National Library of Latvia
  • Culture Information Systems Centre (CISC)

Zinātniskās darbības rezultāts: Nodaļa grāmatā/enciklopēdijā/konferences krājumāKonferences zinātniskais rakstsPētniecībakoleģiāli recenzēts

16 Atsauces (Scopus)

Kopsavilkums

LNCC is a diverse collection of Latvian language corpora representing both written and spoken language and is useful for both linguistic research and language modelling. The collection is intended to cover diverse Latvian language use cases and all the important text types and genres (e.g. news, social media, blogs, books, scientific texts, debates, essays, etc.), taking into account both quality and size aspects. To reach this objective, LNCC is a continuous multi-institutional and multi-project effort, supported by the Digital Humanities and Language Technology communities in Latvia. LNCC includes a broad range of Latvian texts from the Latvian National Library, Culture Information Systems Centre, Latvian National News Agency, Latvian Parliament, Latvian web crawl, various Latvian publishers, and from the Latvian language corpora created by Institute of Mathematics and Computer Science and its partners, including spoken language corpora. All corpora of LNCC are re-annotated with a uniform morpho-syntactic annotation scheme which enables federated search and consistent linguistics analysis in all the LNCC corpora, as well as facilitates to select and mix various corpora for pre-training large Latvian language models like BERT and GPT.

OriģinālvalodaAngļu
Rīkotāja publikācijas nosaukums2022 Language Resources and Evaluation Conference, LREC 2022
Lapas5123-5129
Lapu skaits7
ISBN (Elektroniski)9791095546726
Publikācijas statussPublicēts - 2022
Ārēji publicēts

OECD Zinātnes nozare

  • 1.2 Datorzinātne un informātika

Nospiedums

Uzziniet vairāk par pētniecības tēmām “Latvian National Corpora Collection - Korpuss.lv”. Kopā tie veido unikālu nospiedumu.

Citēt šo