Conference item icon

Conference item

PARSEME 2.0 multilingual corpus of multiword expressions

Abstract:
We present edition 2.0 of the PARSEME multilingual corpus annotated for multiword expressions (MWEs), resulting from efforts of the PARSEME community towards universality-driven modeling of idiomaticity. With respect to previous editions, we extend the annotation scope to all syntactic MWE categories: verbal, nominal, adjectival, adverbial and functional. We cover 17 languages, of which 7 are new. The annotation process is based on cross-lingually unified guidelines, phrased as decision diagrams over linguistic tests, and a typology of 18 MWE categories. The corpus contains almost 5 million tokens, over 250,000 sentences and 140,000 MWE annotations. The applicability of the corpus is tested in baseline experiments with a prompt-based MWE identification system. Results show that generic large language models do not encode sufficient knowledge to solve the MWE identification task.
Publication status:
Published
Peer review status:
Peer reviewed

Actions

Access Document

Files:
Publisher copy:
10.63317/2iy5qf38yhay

Authors


Publisher:
European Language Resources Association (ELRA)
Host title:
Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026)
Pages:
4819-4834
Publication date:
2026-04-30
Event title:
15th Language Resources and Evaluation Conference (LREC 2026)
Event location:
Palma, Mallorca, Spain
Event website:
https://lrec2026.info/
Event start date:
2026-05-11
Event end date:
2026-05-16
DOI:
EISSN:
2522-2686
ISBN:
9782493814494


Language:
English
Keywords:
Pubs id:
2455669
Local pid:
pubs:2455669
Deposit date:
2026-09-10
ARK identifier:

Terms of use


Views and Downloads

Views and downloads will return soon






If you are the owner of this record, you can report an update to it here: Report update to this record

TO TOP