Finding Multiwords of More Than Two Words

Kilgarriff, Adam; Rychlý,  Pavel; Kovář,  Vojtěch; Baisa,  Vít

Informace o publikaci

Finding Multiwords of More Than Two Words

Autoři	KILGARRIFF Adam RYCHLÝ Pavel KOVÁŘ Vojtěch BAISA Vít
Rok publikování	2012
Druh	Článek ve sborníku
Konference	Proceedings of the 15th EURALEX International Congress
Fakulta / Pracoviště MU	Fakulta informatiky
Citace
Obor	Jazykověda
Klíčová slova	collocations; multiword expressions; multiwords; corpus lexicography; word sketches
Popis	The prospects for automatically identifying two-word multiwords in corpora have been explored in depth, and there are now well-established methods in widespread use. (We use ‘multiwords’ to include collocations, colligations, idioms and set phrases etc.) But many multiwords are of more than two words and research for items of three and more words has been less successful. We present three complementary strategies, all implemented and available in the Sketch Engine. The first, ‘multiword sketches’, starts from the word sketch for a word and lets a user click on a collocate to see the third words that go with the node and collocate. In the word sketch for take, one collocate is care. We can click on that to find ensure, avoid: take care to ensure, take care to avoid. The second, ‘commonest match’, will find these full expressions, including the to. We look at all the examples of a collocation (represented as a pair/triple of lemmas plus grammatical relation(s)) and find the commonest forms and order of the lemmas, plus any other words typically found in that same collocation. For baby and bathwater we find throw the baby out with the bathwater. The third, ‘multi level tokenization’, allows intelligent handling of items like in front of, which are, arguably, best treated as a single token, so lets us find its collocates: mirror, camera, crowd. While the methods have been tested and exemplified with English, we believe they will work well for many languages.
Související projekty:	Pattern Recognition-based Statistically Enhanced MT Temporální aspekty znalostí a informací Projekt LINDAT-Clarin - Vybudování a provoz českého uzlu pan-evropské infrastruktury pro výzkum

Jak na přijímačky

Důležité termíny

Přečtěte si o výzkumu na MU

Jak na přijímačky

Důležité termíny

Přečtěte si o výzkumu na MU

Finding Multiwords of More Than Two Words