Corpus Linguist (with Polish, English and Russian) | Scrum Master (PSM I) | Translation & formulaic language scholar | Board member of ESP & ACL journal
We are delighted that our discussions with Aleksandar Trjkla have eventually resulted in a new (and maybe even novel:) theoretical reading of semantic prosody, grounded in the philosophy of language, now available in "Journal of Pragmatics": https://t.co/ICtYgwn8G9
I am honoured to join the Advisory Board of the Czech National Corpus (CNC) research infrastructure for the term 2024-2026. I will do my utmost to provide valuable service to the CNC community, together with other international experts. https://t.co/DeFs5abI1G
Thrilled to finally see our paper on formulaic language in Old English prose in "Journal of Historical Pragmatics". Ania and Piotr, it was pleasure to collaborate with you on that fascinating journey into the past: https://t.co/BBj4a1dgXH
@HJ_Schmid @AusensiJ It often happens when the study is novel and does not fit into recent fashions and fads. Ask for justification. It is this pain of academia that your professional (and also life choices) often depend on some anonymous reviewers' subjective judgements but here: ask the Editor.
(PS.) A follow-up to the previous message. Our preliminary evaluation was conducted using logistic regression and J48 classification; it showed that with circa 0.84 precision and 0.65 recall we can obtain a combinatorial dictionary with thousands of recurrent phrasal equivalents.
(2/2) We used multilingual sentence embeddings and mainly open LLMs: https://t.co/gvJH3UcWim and https://t.co/xNDgvWouO9 as prompt-based classifiers. Preliminary evaluation showed that we can obtain a combinatorial dictionary with thousands of recurrent phrasal equivalents. Ctd
(1/2) It was a real pleasure to share our research at the 4th PACOR symposium on parallel corpora in the picturesque León, Spain. Piotr Pęzik and I talked about extending the English-Polish corpus Paralela and extracting bilingual phraseologies from near-parallel corpus data.
I was honoured to present the results of my recent research and share my reflections on corpus linguists' skills and toolkit in the 2020s at Slovko 2023 in Bratislava. Big thanks to the organizers for an invitation, hospitality and friendly atmosphere! https://t.co/270i7GpEtp
@antlabjp Yes, summarizing patterns of linguistic data sounds very promising, also for lexicographers (notably for languages other than English in the long term). I've found this research particularly interesting: https://t.co/ThBsrMsCht
It takes two to tango, so to speak:) We found that phraseology markers may also break canonical phraseologies or introduce idiosyncratic phrasings unattested or rarely used in native texts. More on this in our new paper available in open access: https://t.co/vqOut2aEde