Contents
Алексеев Д. А., Сидельцев А. В.
О двух конструкциях с вопросительными местоимениями
Алхазова А. В., Дорофеева Е. П., Сатина М. И., Лесняк К. К.
Алхазова А. В., Заруднева А. А., Крайнова А. В., Савельева А. Б., Панич М. Б., Федорова О. В.
Исследования разрешения синтаксической неоднозначности: в поиске оптимальной методики
Auseichyk Yu. V.
Quantitative Profiling and Cluster Analysis of Coordinators in Contemporary French
BБаранов А. Н.
Сфера действия частиц да и нет в ответе на общий вопрос (к феномену «слабой семантики»)
Богуславский И. М.
CChuprina A.
Positional alignment of consonant cluster variability in early speech including effect of morphology
DДобровольский Д. О., Зализняк Анна А., Левонтина И. Б.
Семантические источники показателей чужой речи (на материале немецко-русского параллельного корпуса)
Dolgushin M., Dvoynikova A., Karpov A.
Dragomirov D., Blekanov I., Kulin N., Muravyov S.
RIDER: Reproducible Evolutionary Prompt Search for NLP with Adaptive Operator Allocation
Дубинина Е. С., Котов А. А., Зинина А. А., Аринкин Н. А.
EЕфремова Н.
GGofman S., Lebedev E., Perfilev A., Plohuta-Plakutina E., Shlyakhtina L.
Gruntov I. A.
Semantic profile of Turkish: macroshifts
IIaroshenko P. V., Petrova A. A., Loukachevitch N. V.
Morphology and Emotion Recognition: How Verb Form Affects Language Model Predictions
Инькова О. Ю.
О так называемом изъяснительном значении союза пока
Ionov T., Malykh V.
WorkBench-RU: Can Agents Solve Multi-Step Workplace Tasks Beyond English?
KKalinowski K., Malykh V.
Towards Neural PIE Reconstruction from Heterogeneous Cognate Data
Katsnelson A. I., Lyashevskaya O. N.
Standard and Dagestani Russian in automatic speech recognition: weaknesses of ASR systems
Kharlamova D., Yakuboy A., Vasilev K., Budennaya E.
Can AI Agents Do Diachronic Linguistics? LLMs Meet the Russian National Corpus
Киосе М., Ржешевская А., Потехин В.
Klokova K., Krongauz M., Shulginov V., Yudina T.
Князев С. В., Монина В. Е.
Весь контур — в один слог: интонация А-переспросов в среднерусских говорах
Kolmogorova A., Yavshits E., Majorova M., Dmitrijeva S.
To Shorten or to Lengthen? Output Strategies of ASR Models in Atypical Speech Recognition
Колтунцева В., Алексеева С.
Kopylova E. V., Tsegoeva O. G., Berlin V. A., Vyrenkova A. S., Kuvshinskaya Y. M.
Correcting or Rewriting? An Expert Evaluation of LLM-Based GEC on Academic Learner Data
Коротаев Н. А.
Структура диалогического взаимодействия при совместном построении
Kotelnikova A., Byzov V., Dolzhenkova M., Kotelnikov E.
Can LLM Teams Play What? Where? When?
Kozlova E., Petrova M.,, Baiuk A.
Contradiction Detection in Administrative Offense Rulings: A Case Study Analysis
Krasnobaeva V.
Automatic Processing of Russian Verb-Noun Collocations Using Large Language Models
Kravchenko T. M., Sidorova E. A.
Кравчук М. С., Омельченко М. А., Тыщишина Т. Е., Федорова О. В.
Паузация в речи пациентов с расстройствами шизофренического спектра
LLelik V., Dyachkova М., Dorofeeva S., Lopukhina A., Dragoy O.,, Sekerina I. A.
Limonova A. A., Studenikina K. A.
RuSafety: A Russian Dataset for Evaluating the Safety of Large Language Models
Ли Ю., Вэнь С.
Lukyanchikova A., Ipatova O., Zhilyaeva V., Kolmogorova A., Migal A., Mikhaylovskiy N.
Legal Markup for Reasoning: discursive analysis of tax law court decision texts
Luraghi S., Zanchi C., Brigada Villa L.
Toward a multilingual FrameNet of ancient Indo-European Languages
MМахмутова А., Зиннатуллина А.
Семантическая эволюция русского пейоратива дурачок: корпусное исследование
Marchenko A.
Масленикова А. С., Попова Т. И.
Morozov D.,, Shcherbakova O., Astapenka L.
BelarusianWord Forms Morpheme Segmentation: Data and Evaluation
NNekessova A., Sidorova E., Lamasheva Zh., Sambetbayeva M., Yerimbetova A.,, Kaldarova M., Nazymkhan A.
KazFakeCorpus: A bilingual multi-level corpus for semantic detection of fake news
Николаева Ю. В., Кутович К. А., Колесникова А. Б.
Форма, место или движение? Как устроена иконичность в русском жестовом языке
OObridko M. S.
Automatic Sarcasm Detection in Russian Texts
Orlova E., Bykova V., Kornilov A., Shavrina T.
Scaling Zero-shot Machine Translation on Low-Resource Languages with Linguistic Grammars
PПересыпкина К. А.
Подлесская В. И.
«И лицо, и одежда, и душа, и мысли»: просодия именного сочинения по данным корпусов устной речи
SSadkovskii F.,, Nasyrova R., Sorokin A.
RusHallu-RAG: benchmarking hallucination detection for Russian RAG
Serbina A. P., Borisova V. A., Rabinovich I. A.
Argument Mining for Film Reviews
Shatskikh A., Sorokin A.
Ossetic-COT: Designing a morphologically annotated corpus and morphological analyzer for Ossetic
Shcherbak N., Movsesian A.
Russian Neural Morphological Tagging: Does Grammeme Order Matter?
Shcherbatov D., Lyashevskaya O.
Analyzing Automatic Frame Detection in Natural Language Texts
Shuvalov R. D., Ilina D. V., Kononenko I. S., Sidorova E. A.
A Study of Coreference Resolution for Russian-Language Texts from a Limited Subject Domain
Слюсарь Н. А., Антропова Д. В.
Smirnitskaya A. A.
Solovyev V. D., Ivleva A. I.
Стельмак Е. А.
TTikhomirov M., Zlunikin E., Grigoriev D., Chernyshev D.
LLMTF: An Open Framework and Benchmark for Evaluating Large Language Models in Russian
Timchenko D., Movsesian A.
Extending Retag to Detect Conversion Errors in Syntactic Annotation: SynTagRus as a Case Study
Tiskin D. B.
Russian Pronouns with Focus Antecedents: Coreference and Binding in Corpora
VВасильева Е. В.
Корпусно-статистическая модель оценки укорененности дериватов
YЯнко Т. Е.
Дискурсивное слово ага как маркер пропозициональной установки: функции и просодия
ZЗализняк Анна А.
Типы «отношений смежности» в Базе данных семантических переходов в языках мира
Циммерлинг А. В.
Алексеев Д. А., Сидельцев А. В.
О двух конструкциях с вопросительными местоимениями
This paper compares and contrasts two seemingly closely related wh-indeterminate constructions – the multiple partitive construction (MPC) and the bipronominal distributive construction (BDC). By analyzing data from Russian, Ossetic, and several other languages we arrive at a diachronic dependency between the BDC and a syntactically similar construction, multiple wh-free relatives: a language possessing multiple wh-free relatives also possesses the BDC.
Алхазова А. В., Дорофеева Е. П., Сатина М. И., Лесняк К. К.
Один в поле не воин: статический и динамический подходы к описанию вокализма малых языков (на материале кумыкского и мансийского языков)
The research is devoted to a comparison of static and dynamic approaches to vowel analysis in instrumental phonetics. The static approach, relying on a single spectral section, often fails to reflect the full picture of articulatory processes. The aim of this work is to test, using quantitative metrics, the hypothesis that the dynamic approach yields more accurate data and overcomes the limitations of static analysis. Using material from the Terek dialect of the Kumyk language and the Sosva dialect of the Mansi language, the empirical shortcomings of the static approach are demonstrated, and the advantages of a dynamic description of vowel systems are shown. The results confirm the effectiveness of multi-slice (multiple timepoint) analysis for typologically different languages.
Алхазова А. В., Заруднева А. А., Крайнова А. В., Савельева А. Б., Панич М. Б., Федорова О. В.
Исследования разрешения синтаксической неоднозначности: в поиске оптимальной методики
The study of the syntactic ambiguity of sentences with complex noun phrases with two nouns, for example, A criminal shot the maid of the actress that was hiding him in a closet in the attic of the mansion is of great importance for modeling the process of sentence comprehension in psycholinguistics. Using the traditional verbal questionnaire technique, researchers obtain results that strongly depend on the participants’ individual semantic preferences, which obscures their actual syntactic preferences. In this paper, we used visual context to reduce this effect. The results showed a shift in the participants’ choice from the usual for the Russian «early» closure (‘the maid was hiding the criminal’) towards the «late» closure (‘the actress was hiding the criminal’), which resulted in the absence of statistically significant preferences. The data obtained make it advisable to conduct further research in this direction.
Auseichyk Yu. V.
Quantitative Profiling and Cluster Analysis of Coordinators in Contemporary French
This study presents a quantitative methodology for constructing functional profiles of French coordinators (e.g., et, mais, donc, puis) based on their syntactic properties. Using a representative sample from the Frantext corpus (22.7 million words, 2000–present) and applying copulative (Icop) and connective (Icon) indices, we analyses 56 units via k‑means clustering. The optimal number of clusters (k = 4) is determined by the elbow method and silhouette score, yielding four functional zones: nuclear (4 units), near‑nuclear (7), near periphery (26), and distant periphery (19), which validate a core‑periphery model. Among the genuinely conjunctive uses of the units under study, the copulative and connective functions are almost equally represented; however, in the majority of contexts the units display other types of usage (adverb, discourse marker, subordinating conjunction, etc.). No direct correlation is found between frequency rank and conjunctive potential for most units outside the nuclear zone. The study offers a statistically grounded typology of French coordinators and provides a quantitative assessment of their conjunctionalisation degree.
BБаранов А. Н.
Сфера действия частиц да и нет в ответе на общий вопрос (к феномену «слабой семантики»)
At first glance, the particles da ‘yes’ and net ‘no’ appear to be antonyms. However, the way these units function in answering a yes-no question reveals that this is not the case. For example, the particle da ‘yes’ in single use is not used in answers to simple negative questions (yes-no questions with explicit particle ne ‘no’). The use of the particle da ‘yes’ in answers to ne-li-questions is significantly restricted. There are other specific features of the particle da ‘yes’ as response particle. At the same time, the particle net ‘no’ in single use in such contexts is entirely acceptable. I explain this behavior of the particles da ‘yes’ and net ‘no’ as follows: the semantic scope of the particle da ‘yes’ extends both to the propositional component of the general question (‘P or not P’) and to its attitudinal component the questioner’s assumptions about a possible or preferred answer. Restrictions on the compatibility of the particles da ‘yes’ and net ‘no’ are mostly manifested not in direct prohibitions, but in tendencies that can be traced across large corpora of texts and require the use of corpus technologies. I call this area of semantics “weak interaction semantics.”
Богуславский И. М.
Количественная квантификация и внутренняя сфера действия (к анализу глаголов повторности в русском языке)
Repetitive verbs in Russian (such as povtoryat’ ‘to repeat’, perechityvat’ ‘to reread’, peresprashivat’ ‘to ask again’, etc.) carry a presuppositional component of precedence in their meaning: the current act is related to a previous act of the same type. Ya povtoril ob”yasnenie (‘I repeated the explanation’) presupposes that the explanation had already been given before. Verbs of this group freely combine with quantitative frequency modifiers: povtoril dva raza (‘repeated twice’), povtoril v tretiy raz (‘repeated for the third time’), dvazhdy perechital pis’mo (‘read the letter twice over’). In such combinations, syntactic structure and semantic interpretation may diverge. The syntactic relationship between the adverbials n raz / v n-y raz (‘n times / for the n-th time’) and the repetitive verb suggests that there are n acts of repetition, whereas the meaning of the sentence may explicitly imply that the number of repetitions is one less. The number n characterizes not the repetitions as an independent type of event, but the acts of the base action (of telling, explaining, etc.). For example: Slushayte vnimatel’no, ya dva raza povtoryat’ ob”yasnenie ne budu (‘Listen carefully, I will not repeat the explanation twice ‘). This refers to two acts of explaining but only one act of repetition. The analysis proposed here shows that the key to resolving this problem lies neither in changing the presuppositional status nor in introducing lexical polysemy, but in recognizing the possibility of an internal scope of the quantitative modifier relative to the lexical meaning of the verb.
CChuprina A.
Positional alignment of consonant cluster variability in early speech including effect of morphology
This study explores morphological and positional variation in the production of consonant clusters in early development. An analysis of child speech transcriptions and their adultlike idealised variants is provided in English and French to investigate cluster diversity. A measure of entropy is used to calculate how stable the production of clusters is across word positions (initial, intervocalic, final) and morphological conditions (morphological vs. non-morphological). Results reveal similar positional asymmetries between the child, or a constrained system, and the adultlike, or articulatorily unconstrained variant. Furthermore, morphological complexity destabilises the phonological system at the global scale, illustrated by higher entropy in the final morphological position in English but not French. Finally, gradual convergence between the two systems seems to be prompted by an increase in syntactic complexity of the child speech. Overall, the results are consistent with the statistical learning account of language acquisition.
DДобровольский Д. О., Зализняк Анна А., Левонтина И. Б.
Семантические источники показателей чужой речи (на материале немецко-русского параллельного корпуса)
The article examines the semantic sources of reported speech markers as indicators of one type of indirect evidentiality. Based on data from Russian and German, the following classes of such sources are identified: verbs of speech, demonstrative words, markers of approximate nomination, markers of comparison, modal words, words expressing an attitude towards the quoted utterance and towards the interlocutor, as well as means of parodying reported speech. Furthermore, it is noted that the semantic development of reported speech markers is largely unique; two cases of non-trivial semantic shifts are analyzed. It is shown that German clearly tends to express reportative evidentiality through the conjunctive I and modal verbs, while resorting to words like particles and adverbs in cases where something additional needs to be conveyed: a particular stylistic coloring or a shade of epistemic modality. In Russian, modal verbs are not used in this function; instead, there is a large and ever-expanding class of lexical markers of reported speech—primarily particles and adverbs.
Dolgushin M., Dvoynikova A., Karpov A.
Bimodal detection of cognitive impairment in humans through analysis of speech and textual data in multiple languages
This paper investigates methods for automatic classification of cognitive impairment (patient/healthy) from spontaneous speech in Russian, English, and Chinese. The study uses a preliminary version of the Russian RUBICON corpus and the multilingual TAUKADIAL dataset. A unified pipeline for audio processing (noise suppression, interviewer speech removal) and speech recognition is proposed, using Whisper with few-shot prompts to preserve speech disfluencies. Experiments are conducted using extracted prosodic, acoustic, and linguistic features, as well as neural embeddings (XLM-RoBERTa, ruRoBERTa). The best performance in the audio modality is achieved with prosodic features (UAR 69.13%), while in the text modality the best performance was obtained with embeddings from the fine-tuned XLM-RoBERTa model (UAR 64.35%). Bimodal fusion further improves the best overall result (UAR 70.35%), supporting the presence of cross-lingual markers of cognitive decline.
Dragomirov D., Blekanov I., Kulin N., Muravyov S.
RIDER: Reproducible Evolutionary Prompt Search for NLP with Adaptive Operator Allocation
WepresentRIDER(ReflectiveIterativeDiversity-EnhancedReasoning),ablack-boxframeworkforautomatic promptoptimization. RIDERevolvesapopulationofnatural-languageinstructionswithninepromptoperators, allocatessearchbudgetbyWindowedThompsonSampling,andfeedscurrenterrorcasesbackintomutation.The evaluationisdesignedforreproduciblesame-modelcomparison:baselinesusethesametaskmodel,datapartition, evaluator,andcomparablesearchbudget;finalscoresarecomputedonfullheld-outtestpartitions;themainresultis thesinglebestpromptratherthananensemble;BERTScoreusesmicrosoft/deberta-xlarge-mnli;CommonGenusesmulti-referencescoring;andthepipelineiscoveredby1069automatedtests.OnsixNLPbenchmarks andfivemodel families,RIDERobtains23wins, 0losses, and0tiesagainstZeroShot,APE,EvoPrompt-GA, EvoPrompt-DE,andPromptBreeder.Theaveragemarginoverthestrongestbaselineis7.2percentagepoints.On CodeSearchNet,RIDERobtainsstableBERTScorevaluesof0.769–0.787.
Дубинина Е. С., Котов А. А., Зинина А. А., Аринкин Н. А.
Робот-компаньон в обучении китайскому языку: экспериментальная оценка эффективности и роль невербальных сигналов
This study evaluates the effectiveness of learning Chinese using an anthropomorphic robot companion compared to instruction delivered by teachers with democratic or strict academic interaction styles. The experiment involved novice learners with no prior experience in Chinese across three conditions. Both objective learning outcomes (phonetics, vocabulary, phrase structures, and locative constructions) and subjective evaluations (perceived learning gain, perceived structure, motivation, comfort, and embarrassment) were assessed. In the first stage, no significant differences were found between the robot and the democratic-style teacher in objective performance (p > 0.05). However, the robot group reported higher subjective ratings of perceived learning, instructional structure, and motivation (p = 0.046–0.040). In the second stage, the robot group showed statistically higher results compared to the strict-style teacher across all learning modules (p < 0.01–0.001). At the same time, differences between the two teachers were not significant across all tasks. Subjective evaluations indicated lower levels of embarrassment and higher comfort when interacting with both the robot and the democratic-style teacher (p < 0.05). The obtained effect sizes (r = 0.48–0.82) correspond to medium and large effects. Overall, the results suggest that more structured and predictable interaction in the robotic condition may be associated with better learning outcomes and more positive learning experience. In addition, a possible role of implicit non-verbal cues accompanying robot-based instructions should be noted.
EЕфремова Н.
Корпусный анализ мультимодального конструирования пространства в поведении носителей сокотрийского языка
This paper examines the peculiarities of using the surrounding space in the gestural behavior of speakers of the Socotri language (Socotra, Yemen), as well as the specific features of conveying direction in the construction of narrative space and discourse structuring. The research data consists of video recordings of retellings of the “Pear Film” by 10 L1 speakers of the Socotri language. The analysis revealed such features of gesture as the predominance of using the lower peripheral space in front of the speaker and the lack of distinct changes in the gesture production point. Among the features of conveying direction, a tendency for right-to-left movement can be highlighted, regardless of the type of space being described. We assume that this tendency may, to a certain extent, be produced by the direction of writing in Arabic, which is actively used by the study participants, even though the Socotri language itself does not use any writing system.
GGofman S., Lebedev E., Perfilev A., Plohuta-Plakutina E., Shlyakhtina L.
Building Chukchi Corpora: A Chukchi–Russian Parallel Corpus and a Monolingual Chukchi Corpus for Low-Resource NLP
Many endangered and under-resourced languages lack publicly available corpora suitable for modern NLP, limiting research and downstream applications such as language documentation, morphological modeling, search, and alignment-based text study. We present a new Chukchi–Russian parallel corpus and a complementary monolingual Chukchi corpus compiled from heterogeneous online and printed resources. The parallel corpus was assembled via structured extraction, segment-level alignment, cleaning, and semantic diagnostics, yielding 70,508 ChukchiRussian aligned segments (translation units); the majority of these units are lexicographic word/phrase glosses rather than full sentences. The monolingual corpus contains 15,610 Chukchi segments gathered from untranslated and unaligned materials. We describe data sources, normalization rules for Chukchi orthography, automatic and manual quality control procedures, and corpus statistics. The resulting corpora provide a practical foundation for Chukchi computational research and language preservation efforts.
Gruntov I. A.
Semantic profile of Turkish: macroshifts
This paper presents a systematic description of semantic shifts in Turkish based on comprehensive processing of a Turkish-Russian dictionary. To analyze the macrostructure of semantic variation, we introduce the concept of macroshift, a relationship between large semantic domains (taxa). Based on macroshifts, we construct a semantic profile of Turkish that shows the quantitative distribution of shifts between semantic domains. To our knowledge, this is the first attempt to create a complete semantic profile of an individual language based on systematic dictionary processing. The work was conducted within the DatSemShift project — a database of semantic shifts in the world’s languages. We describe the methodology for working with taxa and the principles of constructing weighted semantic profiles. The material consists of NA Baskakov’s Turkish-Russian dictionary (48,000 words), from which 1,645 semantic shifts were extracted. The proposed approach opens up prospects for comparative analysis of semantic profiles across different languages.
IIaroshenko P. V., Petrova A. A., Loukachevitch N. V.
Morphology and Emotion Recognition: How Verb Form Affects Language Model Predictions
Emotion lexicons are an important tool for textual emotion recognition; however, they represent words as lemmas, disregarding morphological variation. This study tests the hypothesis that grammatical form influences emotion classification. For 473 verbs from the Russian Emotion Lexicon, indicative and imperative forms were generated (5,676 verb forms in total). Three encoders (ruBERT-tiny2, ruBERT-base, ruRoBERTa-large), fine-tuned on a corpus compiled from available Russian-language emotion analysis datasets, classified each form. The results demonstrate that past tense feminine forms systematically achieve optimal performance, surpassing the infinitive in F1-score, whilst imperative forms exhibit the poorest performance. Classification pattern analysis revealed a consistent tendency across all models: imperative and second-person present tense forms labelled as Sadness were systematically assigned to Anger. Preliminary experiments with large language models showed similar tendencies with lower morphological sensitivity.
Инькова О. Ю.
О так называемом изъяснительном значении союза пока
The article analyzes cases where the conjunction poka introduces a subordinate clause governed by verbs of expectation. It is generally assumed that in such cases we are dealing not with circumstantial (temporal) but with complement relations. The author draws on data from the Russian National Corpus, showing the frequency distribution of the conjunctions poka, chto, and chtoby with verbs of expectation, and compares the conditions under which these three conjunctions are used. The analysis suggests that one can speak of a kind of semantic agreement, since poka retains its semantics even when combined with verbs of expectation, although it fills the object valency of these verbs. However, what fills this valency is not the described event, as in the case of the conjunctions chto and chtoby, but rather the moment of its realization: zhdat’, poka R = ‘to wait for the moment when R’. Arguments in support of this claim include the impossibility of inverting the main and subordinate clauses (what is possible with complement relations), restrictions on the aspect and tense of the predicate in the subordinate clause, as well as the fact that poka is not used with the verbs zhdat’ and ozhidat’ in their meaning ‘to assume with a sufficiently high degree of certainty that something may happen’ (epistemic expectation), in which they completely lose the semantics of physical waiting.
Ionov T., Malykh V.
WorkBench-RU: Can Agents Solve Multi-Step Workplace Tasks Beyond English?
We present WorkBench-RU, a bilingual extension of the WorkBench benchmark for evaluating tool-calling LLMagents on realistic workplace tasks requiring multi-step planning, function-call sequencing, and actual database execution. We add Russian localization — queries translated from English, localized tool descriptions, and parallel databases — across five business domains plus cross-domain tasks. During localization, we identify and fix 17 bugs in the original benchmark that had fundamentally compromised evaluation reliability. We evaluate 9 models (1.8B–120B parameters) under 4 reasoning configurations and report state-based correctness averaged across 3 independent trials, along with any@3, all@3, and side-effect metrics. Key findings: (1) structured reasoning (ReAct, native thinking) improves models by up to +20pp, with strongly architecture-dependent effects; (2) Qwen3-32B with native thinking reaches 70.1%, comparable to GPT-4.1-mini (68.1%); (3) English–Russian performance gaps vary from 4pp (Russian-native GigaChat3) to 28pp (GPT), suggesting that training language composition may matter more than model scale.
KKalinowski K., Malykh V.
Towards Neural PIE Reconstruction from Heterogeneous Cognate Data
We construct a multi-source Proto-Indo-European (PIE) reconstruction dataset by aggregating IE-CoR, Kaikki/Wiktionary, Köbler, and EtymologyDB, with additional experiments on synthetic augmentation, validation, and phonetic input variants. We formulate reconstruction as sequence-to-sequence prediction from language-tagged cognate sets to PIE proto-forms and evaluate ByT5 together with data, representation, and architecture ablations. ByT5-Base with conservative Epitran-only phoneticization achieves EM = 0.216 on our PIE test set, a 59% relative improvement over the orthographic baseline. The results show that neural PIE reconstruction is feasible as a hypothesis-generation tool, but performance is driven mainly by data quality and representation fidelity rather than by larger models or aggressive augmentation.
Katsnelson A. I., Lyashevskaya O. N.
Standard and Dagestani Russian in automatic speech recognition: weaknesses of ASR systems
We evaluate the quality of automatic speech recognition (ASR) for Standard Russian and its Dagestani regional variety– a contact idiom formed in the multilingual environment of Dagestan. We tested ten off-theshelf ASR configurations in a zero-shot setting (without fine-tuning) representing three architectural approaches (Conformer+RNN-T, Conformer+CTC, and Encoder-Decoder Transformer) on datasets of Standard Russian (GOLOS, RuDevices; approx. 20K and 90K utterances respectively) and Dagestani Russian (DaGRuS; approx. 40K utterances). All models demonstrate a word error rate (WER)degradation of33%−50%whenevaluatedonDagestani Russian, highlighting a severe domain shift problem. Error structure analysis reveals three primary degradation mechanisms: “Russification” (systematic replacement of Dagestani forms with phonetically similar Standard Russian words); deletion of discourse markers and out-of-vocabulary (OOV) terms such as ethnonyms and toponyms; and generative model hallucinations on incomprehensible audio segments. We found that end-to-end architectures suffer less degradation than cascaded ones (33% − 35% vs. 40% − 42%), yet no model achieves acceptable performance (minimum WER ≈ 43%). Finally, we formulate recommendations for designing ASR systems robust to the regional variability of the Russian language.
Kharlamova D., Yakuboy A., Vasilev K., Budennaya E.
Can AI Agents Do Diachronic Linguistics? LLMs Meet the Russian National Corpus
This paper introduces an integration of Large Language Models with the Russian National Corpus via the Model Context Protocol. We developed an MCP server that acts as an intelligent adapter, abstracting the complexity of the RNC API and enabling models to independently formulate and execute lexico-grammatical searches. To evaluate this system, we conducted a case study using the Diachronicon database, tasked three state-of-the-art models with autonomously reconstructing the diachronic evolution of Russian constructions. The study assesses the models’ capacity to identify developmental stages, provide corpus attestations, and classify changes according to a goldstandard markup. The results show that the corpora-enabled models can conduct multi-step diachronic analyses and often propose valid new stages beyond the gold standard, with Claude Opus 4.5 with extended thinking performing best overall, but overall recall remains low and the outputs are limited by overgeneration of spurious stages, frequent typing/conflation errors, and unreliable corpus attestation, making the approach most effective in a human-in-theloop workflow.
Киосе М., Ржешевская А., Потехин В.
Нарратив на русском языке как иностранном: варьирование локальной структуры дискурса у говорящих с разным уровнем языковой компетенции
The study establishes the relations of local discourse structure in the narratives of L2 Russian language students as mediated by the level of language proficiency (B1/B2 and C1/C2). The research data are the recalls (including 2066 elementary discourse units) of two narrative stories, obtained from 20 learners of Russian. The results in the C1/C2 group display the prevalence of communicative relations (mostly Metadiscursives) and the decrease in Speech disfluencies (mostly Hesitations and Repetitions). Rhetorical relations show difference in the use of Elaboration which prevails in the C1/C2 group, while Joint, Interpretation, Restatement, Antithesis, Sequence prevail in the B1/B2 group. The paper contrasts the results with the ones obtained earlier in the descriptive discourse with the same experiment participants. The results show that in the narrative the number of elementary discourse units is 2.46 times higher, additionally, Self-correction displays growth, mostly in the B1/B2 group. The differences in the use of Sequence and Evaluation are revealed, which stem from discourse genres.
Klokova K., Krongauz M., Shulginov V., Yudina T.
Extracting Intentional Boundaries: Evaluating LLM on Greeting and Goodbye Sequences in Russian Dialogue
The automatic identification of greeting and goodbye sequences in dialogue is crucial for pragmatic theory and for developing socially aware conversational agents. This study evaluates three large language models — YandexGPT 5.1 Pro, GigaChat 2 Max, and GPT-5 — on the tasks of (1) classifying Russian dialogues by presence of these sequences and (2) extracting the precise sequence boundaries. Using 216 dialogues from the Russian Multimedia Politeness Corpus, we employ a prompt-based pipeline with majority voting and multi-tiered evaluation. Results reveal a substantial gap between tasks: while GPT-5 achieves strong dialogue-level classification, and YandexGPT and GigaChat attain perfect precision with conservative strategies, sequence extraction performance declines markedly across all models. Error analysis shows that the primary challenge lies in localizing pragmatic boundaries rather than recognizing intent-bearing content. Models systematically over-include contextual nonverbal actions, yet miss non-conventional markers such as physical gestures, address forms, discourse markers, and future-reference implicatures. These findings underscore that for deep dialogue understanding, models require greater sensitivity to discourse structure, indirectness, and the variability of everyday interaction.
Князев С. В., Монина В. Е.
Весь контур — в один слог: интонация А-переспросов в среднерусских говорах
This paper investigates the intonation of yes-no questions, incompleteness, and A-echo questions in Western and Eastern Middle Russian dialects. An examination of the compression strategies applied to the full melodic contour of yes-no questions and incompleteness within monosyllabic A-echo questions across Middle Russian dialects reveals a general tendency common to all varieties under study: the reduction of the second (extra high) tonal plateau and an increase in the steepness of the rising tonal movement, both conditioned by the lack of segmental material. At the same time, substantial differences are observed in the realization of the first tonal plateau, manifested as an increase in its relative duration in the Western dialects. This points to the distinctiveness of the “stepped” tonal accent characteristic of these varieties. A comparative analysis of A-echo question contours across the different dialects demonstrates that statistically significant differences in the timing of the rising pitch accent are found not only between the Eastern and Western varieties but also within the Western group itself: the “westernmost” dialect, located in the Russian-Belarusian borderland, exhibits a contour with the maximal degree of prominence of the first tonal plateau — the primary diagnostic feature of the “stepped” tonal accent. Intonation of A-echo questions in Standard Russian is similar to that found in the Eastern Middle Russian dialect.
Kolmogorova A., Yavshits E., Majorova M., Dmitrijeva S.
To Shorten or to Lengthen? Output Strategies of ASR Models in Atypical Speech Recognition
This study presents the first systematic comparison of six ASR systems on a dataset of Russian aphasic speech. Using the rates of deletions, insertions and substitutions, we identified three distinct strategies employed by models when handling uncertainty in atypical speech recognition: panicked shortening, lexical preservation, and a stable strategy. Our findings underscore the necessity of considering not only standard quality metrics but also models’ behavioral strategies when developing inclusive speech technologies.
Колтунцева В., Алексеева С.
Ошибки в употреблении имен существительных при чтении вслух связных текстов на русском языке школьниками с нарушениями чтения и без таковых: пример корпусного исследования
This paper presents the construction of an annotated corpus of noun reading errors in Russian and an analysis of its distributional characteristics with respect to age and clinical factors (presence vs. absence of dyslexia). The corpus includes oral readings of connected texts by students in grades 3–4 and 8–11 with and without reading impairments (329 participants in total). The data are organized so that each observation corresponds to the reading of a single noun by a single participant, enabling analysis at the level of a specific grammatical category. The corpus comprises 16,072 noun-reading instances. Quantitative analysis was conducted using generalized linear mixed-effects models. In both age groups, reading impairment status was a stable predictor of error probability, whereas school type and nonverbal intelligence did not show statistically significant effects. Qualitative analysis revealed that the distribution of error types has an internal structure: the most frequent errors include pseudoword formations, substitutions of another noun with matching grammatical features, stress misplacement, syllable-by-syllable reading, and prosodic prolongations. The results demonstrate the potential of the corpus for investigating distributional patterns of oral reading and for developing a resource base for further automated analysis of reading impairments.
Kopylova E. V., Tsegoeva O. G., Berlin V. A., Vyrenkova A. S., Kuvshinskaya Y. M.
Correcting or Rewriting? An Expert Evaluation of LLM-Based GEC on Academic Learner Data
This paper investigates how large language models correct complex grammatical errors in Russian academic learner writing. Unlike traditional minimal-edit GEC systems, LLMs often apply generative rewriting strategies that may improve fluency, but risk structural overcorrection and semantic drift. We introduce a new expert benchmark derived from an authentic 3,1M-word learner corpus and construct an evaluation set annotated for error type and complexity. We propose an expert-driven evaluation framework combining quantitative scoring, structural-change analysis, and blind pairwise comparison. Results reveal a consistent minimal-edit vs. generative trade-off across LLMs. This trade-off has direct implications for evaluation, as purely reference-based metrics may underrepresent structural overcorrection and fail to capture differences in correction strategies.
Коротаев Н. А.
Структура диалогического взаимодействия при совместном построении
The paper analyzes two aspects of interaction between participants of a conversation when they produce syntactic constructions together. It uses data from the multichannel corpus “Russian Pear Chats and Stories”. First, it discusses speech and embodied conduct signals used by first participants to invite their interlocutors to build a coherent structure together or, on the contrary, to indicate that they intend to proceed on their own. Three strategies employed by first participants are distinguished: the open strategy, the closed strategy and the neutral strategy. Quantitative results show that less than one third of all cases of co-construction appear when the first participant uses the open strategy (i.e., explicitly invites the second participant to engage in a joint activity), while most frequently the neutral strategy takes place. Second, the paper shows a connection between the above-mentioned strategies and the turn-taking structure created during co-construction. The open strategy is found to be the most compatible with cases where the second participant completes a single elementary discourse unit initiated by the first participant, and the least compatible with cases of turn extensions. Completions of multi-unit turns occupy an intermediate position with regard to this parameter.
Kotelnikova A., Byzov V., Dolzhenkova M., Kotelnikov E.
Can LLM Teams Play What? Where? When?
Large language models (LLMs) remain limited on tasks requiring indirect reasoning, cultural knowledge, and coordinated hypothesis testing. We investigate whether team-based interaction improves LLM performance in What? Where? When? (ChGK), a quiz game designed to reward collective reasoning. We introduce three team strategies: Voting, Silent Team (the captain observes final answers), and Talkative Team (the captain observes both answers and rationales). To minimize data leakage, we evaluate these strategies on a dataset consisting of 572 ChGK questions released in 2025. Using six recent large-scale open models, we show that team-based strategies outperform single-model baselines, yielding gains of up to 20 percentage points in accuracy. The best team achieves 44.23% accuracy, and approaches human team performance on questions with available human statistics. Analysis of inter-model diversity reveals that disagreement strongly predicts lower accuracy, but explanatory communication substantially mitigates performance drops. We further examine captain behavior and find no evidence of self-preference bias; access to peer rationales improves captain judgments. Overall, LLM teams function primarily as answer selection and error-filtering mechanisms rather than generators of novel solutions. Our findings highlight the importance of interaction and suggest adaptive strategies as a promising direction for multi-agent systems.
Kozlova E., Petrova M.,, Baiuk A.
Contradiction Detection in Administrative Offense Rulings: A Case Study Analysis
The paper presents a pilot system for detecting contradictions between court rulings on administrative offenses and the corresponding laws in Russian legal practice. We formulate the task as a legal-domain Natural Language Inference problem and implement a three-stage pipeline combining rule-based preprocessing, retrieval over a structured repository of legal norms, and LLM fine-tuning with LoRA adapter for pairwise contradiction classification. The evaluation on the test document set shows strong performance in the ternary setting (best scores are precision = 0.96, recall = 0.69, F1 = 0.81). The results demonstrate that combining rule-based filtering, retrievalaugmented search, and domain-adapted large language models can effectively surface legally significant inconsistencies in judicial texts. We discuss current limitations, including retrieval quality and restricted error coverage, and outline directions for improving domain adaptation and extending the approach to other types of regulatory documents.
Krasnobaeva V.
Automatic Processing of Russian Verb-Noun Collocations Using Large Language Models
The automated processing of non-compositional language remains a central challenge for large language models (LLMs), particularly in morphologically rich languages such as Russian. We investigate Russian verb-noun collocations as a testbed for evaluating models’ ability to handle intermediate non-compositional constructions. To this end, we introduce a methodologically balanced dataset of 3,264 expressions covering 479 verbs, equally divided between collocations and compositional constructions. The dataset supports three interconnected tasks-binary classification, single-verb paraphrasing, and contextual gap filling-allowing unified evaluation of both recognition and generation capabilities. We evaluate multilingual and Russian LLMs of varying sizes under zero-, one-, and few-shot regimes and observe near-random classification performance and limited robustness in generation for general-purpose models. While scale and language specialization improve results, they remain insufficient without adaptation. Applying parameter-efficient fine-tuning (LoRA) in a multitask setting yields substantial improvements across all tasks, significantly enhancing precision and semantic consistency. The dataset and code are publicly available at https://github.com/Vera-bahval/russian-verb-noun-collocations.
Kravchenko T. M., Sidorova E. A.
Structure-Oriented Approach to Logical Fallacy Detection in Russian Texts Using Agentic Language Models
This paper presents a structure-oriented pipeline for automatic logical fallacy detection in Russianlanguage texts. The method consists of four stages: argument extraction, entity linking, premise masking, and classification via Natural Language Inference (NLI). By replacing named entities with abstract placeholders, the system isolates the logical skeleton of an argument and compares it against known fallacy templates independently of the topic domain.
Кравчук М. С., Омельченко М. А., Тыщишина Т. Е., Федорова О. В.
Паузация в речи пациентов с расстройствами шизофренического спектра
The article is devoted to the peculiarities of pausing in the speech of patients with schizophrenia spectrum disorders (SSD), studied on the basis of the corpus of reportage and retellings of the “Pear Film” by Wallace Chafe. The corpus consists of speech recordings of young people aged 16-25 with diagnoses of schizophrenia or schizotypal disorder, as well as people without mental illness. The structure of the recordings, which contains a reportage and a retelling based on a single stimulus material, allows us to study speech under conditions of different cognitive loads and reasonably compare the obtained data. Speech studies conducted on the material of this corpus are able to detect features that are not only characteristic of SSD in general, but also allow differentiating similar diagnoses. The present study, devoted to pauses in the speech of patients with SSD, showed that the duration of pauses between utterances significantly differs between the control group and the group with schizotypal disorder during reportage. The obtained result allows us to conclude about the influence of the type of task and the health group on pause patterns and opens up prospects for new research based on the material of this corpus.
LLelik V., Dyachkova М., Dorofeeva S., Lopukhina A., Dragoy O., Sekerina I. A.
RusLan-M: Technical Design and Processing Pipeline of a Longitudinal Multimedia Corpus of Russian Child Speech
This paper presents RusLan-M (v.1.0), an open-access longitudinal multimedia corpus of spontaneous early child speech in Russian, designed for corpus-based research. The corpus consists of video recordings of naturalistic childcaregiver interaction, transcribed and annotated in the CHAT format and publicly available in the CHILDES database. We describe the process of corpus creation, data collection, cleaning and anonymization procedures, as well as transcription and annotation principles. The current version comprises 41 hours of recordings and over 35,000 child utterances, enabling the investigation of longitudinal trajectories of lexical growth, morphological development, syntactic complexity, and patterns of child-directed speech in Russian. Future development directions include corpus expansion with additional longitudinal datasets, systematic manual validation of automatic annotation tiers, and further integration of automatic alignment and annotation tools (BatchAlign2) adapted to Russian child speech. RusLan-M is designed for research on language acquisition, corpus linguistics, and the development of computational methods for morphologically rich languages.
Limonova A. A., Studenikina K. A.
RuSafety: A Russian Dataset for Evaluating the Safety of Large Language Models
The widespread deployment of Large Language Models (LLMs) raises critical concerns about their safety, particularly in languages other than English. While recent research demonstrates that LLM security degrades outside English, safety in Russian remains largely underexplored. In this work, we introduce RuSafety, a curated dataset of 1,198 dangerous Russian prompts covering 14 commonly used safety scenarios. The dataset is derived from the multilingual XSafety benchmark through a multi-stage filtering and validation procedure. We evaluate four contemporary LLMs: YandexGPT Lite, YandexGPT Pro, GPT-OSS-20B, and GPT-OSS-120B. For evaluation, we employ an automatic LLM-as-a-Judge framework with Qwen3-235B as the safety assessor. While all models achieve high overall safety rates when tested exclusively on dangerous prompts, they rely on markedly different alignment strategies. Our findings highlight the importance of language-specific safety evaluation and demonstrate that highlevel safety metrics alone are insufficient to capture differences in model alignment strategies.
Ли Ю., Вэнь С.
Корпусное исследование конструкций китайского языка — на примере конструкции негативной оценки в Базе данных CCGD
This paper analyzes 60 negative evaluation constructions in Modern Chinese from the CCGD database along three dimensions: grammatical, semantic, and diachronic. Correlations are identified between slot type and evaluation object, negation explicitness and indirectness, and semantic type and formation mechanism. Over half of the constructions lack explicit negation markers; the form-meaning asymmetry in this class is organized into five ordered types forming a scale of indirectness.
Lukyanchikova A., Ipatova O., Zhilyaeva V., Kolmogorova A., Migal A., Mikhaylovskiy N.
Legal Markup for Reasoning: discursive analysis of tax law court decision texts
Recently, reasoning Large Language Models (LLMs) have emerged as a prominent subject in computational linguistics. However, reasoning models are mostly trained on computationally verifiable STEM data. To address this limitation, we introduce a structured framework of legal reasoning with potential applications in future research. We study a legal text structure using automated discourse annotation and explore the most stable discursive patterns, which constitute logic of legal court decisions texts. Specifically, we find that the common legal logic primarily follows the “cause → legal_ref → conclusion” pattern, and most of the functions are linked by the ‘“cause”, which confirms its structural importance in legal reasoning as the main connector. We strongly believe that this finding will facilitate the construction of reasoning datasets in legal and other non-STEM domains.
Luraghi S., Zanchi C., Brigada Villa L.
Toward a multilingual FrameNet of ancient Indo-European Languages
FrameNet is an English-based lexical database that shows how words are used by providing information as to which participants and relations are evoked by a certain concept. Recent efforts toward a multilingual FrameNet have not targeted either ancient languages or different historical stages of the same language. In our paper we propose creating a multilingual FrameNet for Ancient Indo-European languages starting with a set of 80 verb meanings annotated in the Pavia Verbs Database. Our pilot study includes four verb meanings: RAIN, THUNDER, SEE, LOOK AT. As the adequacy of the semantic frames developed for English turns out not to be appropriate for the languages in our sample, we propose two new frames that can account for the analyzed data.
MМахмутова А., Зиннатуллина А.
Семантическая эволюция русского пейоратива дурачок: корпусное исследование
Drawing on data from the Russian National Corpus (RNC) and social media, this study traces the semantic
evolution and frequency dynamics of the pejorative diminutive durachok, a diminutive of durak ‘fool’, in comparison
with the lexemes durak, dura, durochka, idiot, and idiotka. Normalized frequencies were obtained and a contextual
analysis of 450 occurrences of durachok was carried out for three socio-historical periods (1960-1985, 1986-2000,
and 2001-2020). The results reveal a marked shift from a balanced distribution of affectionate, pejorative, and ironic
uses in the late Soviet period to the predominance of pejorative meanings in the contemporary era. Alongside this
qualitative shift, the overall frequency of durachok steadily declines, in line with the general decrease in the
productivity of diminutives. A collocation analysis conducted with the RNC’s “Word Sketch” (Portret slova) tool
brings to light both the semantic specialization of the lexeme and the persistence of certain phraseological patterns
and cultural archetypes. The incorporation of social media data documents an accelerated pejoration of durachok in
digital communication. Interpreting the findings in terms of pragmaticalization and subjectification [10], the study
contributes to research on the dynamics of evaluative vocabulary in Russian.
Marchenko A.
Acted vs Crowd-Sourced Emotion in Russian Speech: Acoustic Features and 49-Layer Neural Probing on Dusha and RESD
We compare 27 phoniatric acoustic features with layerwise representations from wav2vec2-xls-r-1b (1B parameters, 49 layers) for speech emotion recognition on two Russian corpora: Dusha (crowd-sourced elicited speech,
2,000 samples, 4 emotions) and RESD (studio-acted speech, 790 samples, 4 emotions). Linear probing reveals that
acted speech peaks at layer 3 (F1 = 0.793), while crowd-sourced speech peaks at layer 35 (F1 = 0.769), with a crossover near layer 14. For context, fully fine-tuned HuBERT on a comparable Dusha subset reaches F1 ≈ 0.81 (Kondratenko et al., 2023); the gap with our 0.769 is expected for frozen-backbone probing and locates the upper bound
for representation-only methods. With matched classifiers(LogReg), acoustic features cover 73.5% of neural model
performance for acted speech (F1 = 0.583 vs. 0.793) and 62.5% for crowd-sourced speech (F1 = 0.481 vs. 0.769).
Cross-corpus transfer (Dusha→RESD peak F1 = 0.462 at L36, RESD→Dusha peak F1 = 0.429 at L35) shows
persistent asymmetry: crowd-sourced training generalizes better to acted speech than vice versa (∆ = +0.033). Ablation study shows that MFCC dominates for acted speech (∆F1 = −0.199), while frequency parameters contribute
most for crowd-sourced (∆F1 = −0.052). We interpret the divergence as evidence that acted and elicited emotions
rely on different production mechanisms and consider how these findings relate to Russian data protection laws
(FZ-152, FZ-572).
Масленикова А. С., Попова Т. И.
Сравнение методов автоматической разметки речевых формул в русскоязычном интернетдискурсе: пилотное исследование
This study focuses on developing and comparing methods for automatic annotation of speech formulas in a
corpus of Russian internet comments. Speech formulas are a class of multiword expressions that convey emotional
reactions in dialogue. The research material consisted of a corpus of 10,000 comments (157,261 tokens) collected
from five Telegram channels. Dictionary-based formal search using 437 units achieved 21% precision. To improve
accuracy, three methods were developed: Random Forest classification (56% precision), dependency parsing-based
syntactic filtering (73.3% precision, 8.7% recall), and punctuation-based filtering (76.4% precision, 74.0% recall).
Analysis revealed that syntactic parsers systematically misclassify interjective units: 68.5% of true speech formulas
received the advmod label instead of the correct ROOT. The punctuation-based method showed the best results,
improving precision by 3.64 times over the baseline formal search. The research demonstrates that for linguistic
phenomena with clear formal markers, simple rule-based methods can outperform machine learning, especially with
limited annotated data.
Morozov D.,, Shcherbakova O., Astapenka L.
Belarusian Word Forms Morpheme Segmentation: Data and Evaluation
Morpheme segmentation is a crucial step for morphological analysis and linguistically informed subword tokenization, particularly for highly inflectional languages. However, research on low-resource languages like Belarusian
has been hindered by the lack of word-form datasets, with existing resources limited strictly to dictionary lemmata.
To address this gap, we introduce Slounik-Wordform, the first large-scale dataset of Belarusian word-form morpheme segmentations. It comprises 332,497 expertly validated entries generated semi-automatically from a lemmabased dictionary. Using this novel resource, we benchmark several segmentation architectures, including CNNs,
LSTMs, and various monolingual and multilingual BERT-like models. To robustly assess the models’ generalization capability to out-of-vocabulary morphemes, we evaluate them under three distinct data-splitting scenarios:
random, lemma-based, and root-based. In contrast to prior studies on lemmatized data, our results demonstrate
that fine-tuning large multilingual BERT-like models significantly outperforms traditional neural networks. Specifically, the XLM-RoBERTa-large model achieves state-of-the-art performance, reaching a word-level accuracy of
99.1% on the random split, 92.5% on the lemma-based split, and 77.7% on the root-based split.
Nekessova A., Sidorova E., Lamasheva Zh., Sambetbayeva M., Yerimbetova A.,, Kaldarova M., Nazymkhan A.
KazFakeCorpus: A bilingual multi-level corpus for semantic detection of fake news
The paper addresses the lack of bilingual annotated resources for automatic fake news detection in the KazakhRussian media space. We introduce KazFakeCorpus, a balanced bilingual corpus annotated using a multi-level
semantic scheme that captures message reliability, fake content type, disinformation technique, communicative intent,
modality, and source characteristics. The corpus was constructed from verified Gov.kz materials for the REAL class
and synthetically transformed texts for the FAKE class. After preprocessing and balancing, the final dataset contains
4,276 texts in Kazakh and Russian. Annotation was performed in Label Studio by two independent experts, while a
pilot phase of 120 texts was used to refine annotation categories and guidelines. Inter-annotator agreement measured
with Krippendorff’s Alpha ranged from 0.79 to 0.88, indicating substantial consistency. Corpus analysis revealed the
prevalence of misattribution, clickbait, and emotional pressure techniques. The proposed framework models fake
news as a structured semantic phenomenon and supports research in disinformation analysis, explainable NLP, and
cross-lingual fake news detection for low-resource languages.
Николаева Ю. В., Кутович К. А., Колесникова А. Б.
Форма, место или движение? Как устроена иконичность в русском жестовом языке
The article presents the results of a quantitative study on iconicity in the Russian Sign Language (RSL). Based
on 710 signs from 14 thematic sections of the dictionary “Govoryashchiye Ruki” [5], the analysis focuses on which
formal parameters of a sign – handshape, location, or movement – most frequently carry iconicity, how they combine,
and how iconicity is distributed across different semantic fields. Seventy-one percent of the signs are iconic in at least
one parameter. Movement is the most frequent carrier of iconicity (61.5%), while location is the least frequent (21%).
Iconic handshape and location rarely act alone, but movement alone is the source of iconicity in nearly a third of the
cases. A statistically significant correlation exists between a sign’s thematic category and its iconicity profile: in
“Animals” and “Clothing”, location iconicity is more frequent; in “Sports and Leisure,” movement iconicity is more
frequent; and in “Tableware” and “Household Items,” handshape iconicity is more frequent. Compared to the Italian
Sign Language, RSL shows similarities (around 50% iconicity across two parameters) and differences (a higher
frequency of body location iconicity). The results demonstrate that iconicity is a significant systemic feature of the
Russian Sign Language. The proposed methodology allows us to quantify its occurrences in various types of visual
signs – ranging from gesticulation to RSL signs – and clearly demonstrates its gradual nature.
Obridko M. S.
Automatic Sarcasm Detection in Russian Texts
This article is devoted to sarcasm detection in short written texts. While the majority of previous works focused
on English data mainly using Twitter and Reddit posts as a resource, this research explores a new language (Russian)
and a new speech genre of negative reviews that lacks any explicit sarcasm marking. 3596 reviews were scraped from
otzovik.ru and pravogolosa.net. The collected data was manually assessed by three linguists, showing ambiguity of
sarcasm comprehension even by native speakers. All assembled data is available online. The collected texts were used
for fine-tuning a RuBERT model, previously finetuned for a sentiment detection task. The final sarcasm-RuBERT
model surpassed results shown in similar researches on Russian data with a F1-score of 85.96 and performed better
than GPT-4. In addition to annotated corpora and fine-tuned model sarcasm was theoretically examined. Both the
fine-tuned model, GPT-4 model and assessors pointed out similar relatively definite sarcasm markers such as quotation marks, parentheses, words with a high sentiment and others. More ambiguous sarcasm types such as words with
a “local” positive sentiment, rhetorical questions, wishes and other were also theoretically researched, all contradicting
an explicit sentiment of the utterance and heavily relying on the context.
Orlova E., Bykova V., Kornilov A., Shavrina T.
Scaling Zero-shot Machine Translation on Low-Resource Languages with Linguistic Grammars
Most of the world’s languages lack parallel corpora, monolingual web data, or sufficient representation in
multilingual LLM pretraining. However, many are documented in descriptive grammars containing exhaustive
syntactic information, interlinear glossed examples and translations. Recent work has shown that large language
models can leverage grammars in-context for zero-shot translation and typological classification, but it remains
unclear whether grammars alone can serve as the primary supervision for training machine translation systems.
We introduce a scalable grammar-centered framework for machine translation, converting descriptive grammars
into structured LLM-readable context by extracting example sentences, glosses, and leveraging available external
typological metadata. Across five typologically diverse low-resource languages, we present structure-aware regimes that incorporate gloss-level information and grammar conditioning to enable in-context learning for machine
translation. We further analyze how translation quality scales with the number of grammar-derived examples –
for Georgian (49.01 ChF), Chamorro (53.32 ChF), Basque (63.90 ChF), Igbo (52.83), and Korean (37.09 ChF)
languages. Our results show that (1) grammars alone can support non-trivial translation performance when nothing
else is available, (2) incorporating typological metadata does not show consistent generalization over flat sentence
pairs, and (3) sentence-level examples matter most: extracted parallel examples consistently improve translation
quality relative to zero-shot and typology-only prompting.
This work reframes descriptive grammars as computational resources rather than static references, offering a
practical pathway toward scaling machine translation to languages that have no corpus data. The code for the
project will be publicly available at: https://github.com/MTOB-HSE
Пересыпкина К. А.
Коммуникативная матрица фотографического инскрипта: гендерный аспект (по данным архива «История России в фотографиях»)
This paper aims to establish the dependence of quantitative and structural parameters of photographic inscriptions
on the numerical composition and gender (binary model) of communication participants (giver/receiver positions)
as well as to analyze the communicative matrix focusing on its four types: “male → male”, “male → female”,
“female → female”, “female → male”. The research material comprises 546 inscriptions extracted from the
photograph metadata of the “History of Russia in Photographs” collection (filtered search by the keyword “надпись”
(‘inscription’), automated data collection, manual annotation, and subsequent verification by comparison with
the rectos of photographs, which confirmed the reliability of the annotation): Другу Леониду от Гришкина Геннадия
14/III – 1967 год (To friend Leonid from Grishkin Gennady 14/III – 1967). The findings reveal that the quantitative
and gender factors determining the communicative roles of addresser and addressee influence not the content
of the inscription (the length and structure of inscriptions do not depend on the gender of communicants), but
the practice of gift-giving — its frequency and directionality. The study identifies the predominance of individual
communication, oppositely directed asymmetry of communicative roles (men act as givers significantly more often,
women as receivers), the dominance of intragender addressivity, especially pronounced among women, and other
specific features.
Подлесская В. И.
«И лицо, и одежда, и душа, и мысли»: просодия именного сочинения по данным корпусов устной речи
The study investigates prosodic shaping of coordinated noun phrases. The material analyzed comes from the
Russian National Corpus (RNC), specifically the Multimedia subcorpus and the Multipark subcorpus, as well as from
the pilot version of the corpus of spontaneous personal narratives “What I Saw.” Two principal prosodic strategies of
nominal coordination have been identified: the adaptive and the parallel strategies. Under the adaptive strategy, the
group is prosodically integrated: the last conjunct serves as the prosodic representative of the entire group and receives
prosodic shaping determined by external factors – namely, the group’s position within the overall communicativeprosodic structure. The preceding conjuncts are prosodically dependent on the final one – either unaccented, or their
tonal movement is constrained, in particular, its direction being mirror-opposite to that of the final conjunct. The
adaptive strategy may appear in the flat version, where there is a single prosodic center in the chain, or in a branching
version, in which subgroups with more closely interrelated members may be embedded within the group. Under the
parallel strategy, the conjuncts are prosodically autonomous: each bears an accent, their tonal movements are aligned
and unidirectional, and each receives prosodic realization determined by external conditions – that is, by the group’s
position within the overall communicative-prosodic structure. It is shown that the mechanism of prosodic adaptation
functions as one of the means of maintaining discourse coherence, alongside syntactic and referential devices.
Sadkovskii F.,, Nasyrova R., Sorokin A.
RusHallu-RAG: benchmarking hallucination detection for Russian RAG
This paper introduces RusHallu-RAG, the first comprehensive benchmark for detecting hallucinations in
Retrieval-Augmented Generation (RAG) systems specifically for the Russian language. Addressing a significant
gap in the field, we construct a dataset of 1,000 query-answer pairs, split between general knowledge (SberQuAD)
and scientific domains (ruSciBench). To create a challenging and realistic evaluation environment, we implemented a perturbation strategy on the retrieved documents, controlling for the position of the first relevant document
and thus the difficulty of the information extraction task. A novel, fine-grained taxonomy of six hallucination
types—Contradiction, Unconfirmed, Missing, Excess, Partial, and Oversight—is proposed and used for human
and LLM-based annotation at both the response and span levels. The benchmark is used to evaluate a diverse suite
of 15 open-weight and proprietary LLMs as hallucination detectors. The results reveal that model scale does not
guarantee performance, with medium-sized models (20B-33B) sometimes outperforming larger ones. However, the
proprietary Gemini-2.5-Pro model significantly outperforms all open-weight models across all tasks, highlighting
a current gap in accessible, high-quality Russian-language hallucination detectors. The benchmark and code are
publicly released to foster further research.
Serbina A. P., Borisova V. A., Rabinovich I. A.
Argument Mining for Film Reviews
The paper presents the creation of a new dataset for Argument Mining (AM) in Russian, focusing on film reviews.
The dataset, a reworked version of the blinoff/kinopoisk collection, consists of 3,280 reviews annotated for stance and
argumentative premises. A novel rule-based system is introduced for determining the final argument label, which
accounts for the nuanced balance of pros and cons statements in a review. The study evaluates two transformer-based
models fine-tuned on the dataset: DeepPavlov/rubert-base-cased and ai-forever/ruRoberta-large. The performance
of the models is analyzed in detail, testing hypotheses related to a film’s release period (20th vs. 21st century), its
rating position (top-250 vs. bottom-100), and the congruence between stance and premises. The results demonstrate
that ruRoberta outperforms ruBERT for stance and premise detection. The analysis reveals only one statistically
significant finding: both models perform substantially better on congruent texts, where stance and premise align. The
apparent differences in film century and rating do not reach statistical significance given the current test sample size.
The newly created dataset and the fine-tuning of encoder models establish a foundation for future research in argument
mining for Russian-language texts.
Shatskikh A., Sorokin A.
Ossetic-COT: Designing a morphologically annotated corpus and morphological analyzer for Ossetic
In this work we present the first morphologically annotated corpus for Iron Ossetic that conforms to the Universal
Dependencies schema. The corpus includes 5454 manually annotated sentences from the Iron Ossetic Corpus of
Oral Texts, containing 74032 tokens. We use this corpus to train a BERT-based morphological analyzer. The
analyzer achieves tag accuracy of 95.60%.
Shcherbak N., Movsesian A.
Russian Neural Morphological Tagging: Does Grammeme Order Matter?
The paper addresses the problem of neural morphological tagging for Russian. Existing research demonstrates
that tagging performance improves when a set of morphological features (grammemes) is predicted sequentially,
using each previous prediction to generate the next. Furthermore, studies in other domains show that the order
of output generation can significantly impact model performance. However, the influence of this order remains
unexplored in the context of morphological tagging. This paper investigates whether the performance of neural
morphological tagging depends on the order in which grammemes are predicted. This study implements a sequential grammeme prediction model and conducts experiments with different prediction orders selected from the
literature. We also employ a loss function whose value is order-invariant. The study was conducted on five Russian
language corpora. For most corpora, we observed a statistically significant difference in the overall performance
metric between the best-performing and worst-performing prediction orders. Furthermore, each corpus contained a
set of grammemes whose F-scores differed significantly between these orders. These sets predominantly included
grammemes for nominative and genitive cases, verb aspect, and animacy. The F-score difference for these grammemes reached 0.36 (with an average F-score of 97.14). The results showed that the prediction order affects both the
overall tagging performance and the prediction of individual grammemes. We demonstrated qualitatively how the
interaction of grammemes affects the model’s performance.
Shcherbatov D., Lyashevskaya O.
Analyzing Automatic Frame Detection in Natural Language Texts
Frame resources such as Berkeley FrameNet cover a wide range of stereotyped situations encoded in language.
They also specify the participants and the relationships between them. Typically, the task of annotating text with
frames and identifying semantic gaps in the FrameNet structure is completed by experts. They then create new
frames to fill those gaps. This paper investigates the effectiveness of large language models in the task of automatic
identification of Frame Evoking Elements (FEEs). It was important to identify both FEEs for which there are
corresponding frames and FEEs for which there are no frames. We used a large language model and few-shot
learning to annotate texts with frames in English, Spanish, and Russian. We also manually annotated the same texts
to evaluate the model’s performance. The results reveal a clear performance divide based on the target language
with English being the highest-scoring and Russian and Spanish scoring lower metrics. This performance gap was
attributed to factors such as inherent morphological complexity and greater lexical variation in Spanish and Russian.
The additional analysis of common errors revealed several limitations, such as preference for abstract frames and
hallucinations (generating non-existent frames). Such work could be useful when building a Frame Completion
System, which would annotate Frame Evoking Elements in the given text and then automatically create new frames
when necessary.
Shuvalov R. D., Ilina D. V., Kononenko I. S., Sidorova E. A.
A Study of Coreference Resolution for Russian-Language Texts from a Limited Subject Domain
The paper presents a study of coreference resolution methods in Russian‑language popular science texts of a
limited thematic scope. A distinctive feature of the proposed approach is its treatment of coreference resolution within
the broader context of information extraction and knowledge graph construction. To this end, a specialized dataset
has been created, with annotation principles adjusted to work with entity mentions belonging to a specific subject
domain. In particular, singletons and abstract mentions have been added to the annotation scheme, and each mention
was labeled with a class of the subject domain entity. The dataset‑creation methodology included an initial
LLM‑based (few‑shot) annotation of texts followed by manual correction. Further experiments with the
Qwen2.5‑32B‑Instruct model showed that the LLM performs worse on the task than smaller models trained on the
datasets. The created corpus includes 21 texts on computational linguistics and contains annotations for 9,905
mentions and 2,683 entities (clusters). Experimental studies were based on the SFT approach and involved examining
the transferability of a base model trained on a large universal dataset to domain‑specific texts, as well as the impact
of semantic information about mention classes on the results. The best performance was achieved by the approach in
which for each entity class, a trainable embedding was created, after that concatenated with the vector representation
of the mention, and then, fed to the input of a linear layer, and the model was fine‑tuned on the training portion of the
domain dataset. This approach yielded an F1 of 0.626 for mention linking and an F1 of 0.346 for coreference
resolution.
Слюсарь Н. А., Антропова Д. В.
Роль порядка контролера и мишени в обработке согласования по роду и числу: экспериментальное исследование на материале русского языка
Agreement feature processing is one of the core issues of psycholinguistics. Previous studies mainly focused on
number and gender features and obtained controversial results: in some studies, gender was more salient in processing,
while in others number was more salient. The studies relied on different languages, used different constructions and
different methods. The goal of our study was to explore one of the factors that could potentially explain the observed
discrepancies – the order of controller and target of agreement. We conducted a word-by-word self-paced reading
study comparing gender and number agreement errors in affirmative sentences and questions. These two types of
sentences imply controller-target and target-controller orders respectively. We observed different patterns for
affirmative sentences and questions. In questions, gender errors were costlier than number errors. However, in
affirmative sentences the tendency was the opposite. We proposed two hypothetical explanations for our results and
suggested ways to test them experimentally.
Smirnitskaya A. A.
Is metonymy the prevalent cognitive mechanism of semantic shifts? Evidence from the DatSemShift database
This article examines the DatSemShift database of semantic shifts, reassessing its significance as a resource for
investigating cognitive patterns of semantic change. Drawing on many years of work with the database, the author
identifies and systematises the cognitive mechanisms underlying semantic shifts within the stages of the shift
development. I suggest to identify the four stages of semantic shift: the preliminary state; the stage of the first
cognitive action of the first Speaker; the stage of the first cognitive action of the first Listener (corresponding to the
“invited inference” by E. Traugott) and the stage of the further dissemination (cf. “conventionalization of
implicature”). Further I identify four cognitive mechanisms for the stage of the first Speaker’s activity. These include
the transfer based on external similarity, metaphorical extension, “classical” metonymical extension and the transfer
based on “situational metonymy”. The factor of external motivation and the phenomenon of syncretism are also
considered. Analysis of the 300 shifts in the DatSemShift 3.0 database with the greatest number of realisations, shows
metonymy to be the prevalent cognitive pattern. This is also confirmed in the colexification data in the conceptually
related CLICS database.
Solovyev V. D., Ivleva A. I.
Beyond Accuracy: Systematic Errors in Russian Morphological Analyzers and the Ensemble Consensus-based Approach
While numerous tools for Russian morphological analysis exist, their systematic errors and biases can undermine
downstream linguistic applications. This paper presents a comprehensive evaluation of six state-of-the-art Russian
morphological analyzers (Natasha, SpaCy, pymystem3, DeepPavlov, Stanza, UDPipe). Moving beyond aggregate
accuracy scores, we conduct a fine-grained quantitative and qualitative analysis of systematic errors using the
GRAMEVAL 2020 and SynTagRus datasets. Our analysis reveals significant biases and systematic patterns in
lemmatization and POS-tagging. To mitigate these individual biases, we propose and validate two consensus-based
ensemble methods which achieve statistically significant improvements over the best individual analyzer (from
+0.31% to +2.14% gain). Our results demonstrate that while no single tool is universally optimal, their complementary
strengths can be leveraged to produce more reliable annotations for linguistically sensitive research. Still, no
consensus eliminates errors shared by all or a majority of tools.
Стельмак Е. А.
Перестать верить на слово: новый синтетический датасет для верификации утверждений и тестирования на устойчивость к галлюцинациям у русскоязычных БЯМ
Large language models (LLMs) demonstrate a tendency to generate hallucinations — information that does not
match the data from the target source or cannot be confirmed from the context. Most modern LLMs quality assessment
tools, such as RAGAs, are considered multilingual, but they are significantly worse at handling data in Russian. In
addition, existing methods rely on automatic text generation, without considering the linguistic structure of the texts.
Our method allows us to evaluate the stability of the LLMs to generate and recognize false data and identify the
linguistic causes of their occurrence. Using the developed methodology, typical linguistic constructions that most
often lead to hallucinations are identified and described. The article describes the process of creating and applying a
benchmark based on industrial engineering and linguistic analysis methods. The benchmark was created using data
synthesis: LLMs was used to generate pairs of statements based on Wikipedia texts. Each pair included an original
statement from the article and a semantically similar but contradictory statement. To evaluate the LLMs, pairs of
statements were submitted to the input of a model with a prompt to assess the truth of a deliberately false (contradictory) statement. The computing power of the Groq cloud service was used to conduct experiments. The results of
experiments on Llama-3.1-8b-Instant, ChatGPT-5.5, Qwen3.6-Plus prove the effectiveness of the proposed method:
about 31% of the Llama’s texts and 21% ChatGPT’s texts generated based on the benchmark contain hallucinations.
Seven linguistic patterns contributing to hallucinations have been identified. It has been found that hallucinations are
more often provoked by a combination of patterns rather than individual constructions.
Tikhomirov M., Zlunikin E., Grigoriev D., Chernyshev D.
LLMTF: An Open Framework and Benchmark for Evaluating Large Language Models in Russian
The rapid development of Large Language Models (LLMs) requires reliable and scalable evaluation tools.
However, the research community faces significant fragmentation, outdated frameworks designed for foundational
models, and closed proprietary benchmarks. These problems are particularly noticeable in the Russian-language
segment, which suffers from a lack of high-quality datasets and open tools. To address these challenges, we introduce LLMTF1
, a novel open framework and benchmark for evaluating LLMs. LLMTF unifies the evaluation
process through a message format, supports various backends (vLLM, HuggingFace, OpenAI API), and offers a
comprehensive benchmark comprising 55 datasets across 5 key categories: knowledge, skills, long context, RAG,
and instruction following. Additionally, it features an LLM-as-a-Judge Arena with style control for robust openended evaluation, and integrates specialized benchmarks for complex tasks like extreme long-context summarization and encyclopedic generation. Our framework provides research groups with an accessible and local tool to
effectively evaluate models, identify their strengths and weaknesses, and analyze the impact of changes or model
modification processes on generation quality.
Timchenko D., Movsesian A.
Extending Retag to Detect Conversion Errors in Syntactic Annotation: SynTagRus as a Case Study
The Universal Dependencies (UD) project has grown rapidly through semi-automatic conversion of existing
treebanks, but ensuring the quality of converted annotations remains a challenge. Manual verification does not
scale, and existing automatic methods cannot distinguish conversion errors from inherent annotation complexity.
We present a method that addresses this gap by training two parsers on a token-aligned parallel corpus: one on
the original annotation and one on its UD conversion. By requiring correct predictions from the parser trained on
the original annotation, our approach isolates errors specifically introduced during conversion while filtering out
cases where the construction is simply difficult to parse. We demonstrate the effectiveness of this method on the
SynTagRus corpus and its UD counterpart. To enable a direct comparison, we created a fully token-aligned version
of the two corpora, resolving differences in tokenization and ellipsis representation. We also proposed a simple
method for aligning syntactic relations across the two corpora, addressing the fact that relations involving the same
token do not always correspond due to differences in annotation schemes. Our analysis identified several hundred
errors in the test set. These comprise six distinct types of conversion errors, three of which persist in the current
converter, and ten groups of annotation inconsistencies between the old and new corpus parts. Our method offers
a practical, scalable tool for conversion error detection and is applicable to any language pair with aligned original
and converted annotations.
Tiskin D. B.
Russian Pronouns with Focus Antecedents: Coreference and Binding in Corpora
Despite a lot of interest for the factors influencing the choice of pronoun (reflexive or personal) with an antecedent in Russian, the role of the anaphotic relation—coreference or semantic binding—has been understudied,
including disagreements as to the acceptability of particular data points. To clarify things, I employ large corpora
(Araneum and GICR) to study the role of several factors in determining the likelihood of the pronoun being used as
coreferential or bound: pronoun type (reflexive/personal), pronominal morphology (nominal/adjectival), syntactic
distance to the antecedent as well as word order (cf. (Wurmbrand 2017) for German pronominals). I conclude that
Russian reflexives are capable of, and frequently used with, coreferential interpretation (and more so when preposed
to the antecedent) and that Russian 1st- and 2nd-person pronouns can be used as bound (i.e. as fake indexicals) when
postposed.
Васильева Е. В.
Корпусно-статистическая модель оценки укорененности дериватов
The article proposes a corpus-based statistical approach to assessing the degree of entrenchment of non-codified
derivatives in the modern language. Entrenchment is considered a measurable characteristic, identified on the basis
of observable parameters of a word’s functioning in textual data. The study relies on a specialized corpus of contexts
containing derived units, which includes text fragments (snippets) with meta-information about the functional style,
source, and time of usage.
The proposed model integrates three independent indicators: contextual reproducibility, defined as the number
of unique textual occurrences of the derivative; diachronic evidence, reflecting the temporal span of the derivative
unit’s functioning; and the representation of the word across various functional styles.
The obtained indicators are interpreted as empirical measures of the degree of entrenchment of derivatives in
linguistic practice. The proposed approach demonstrates the possibility of formalizing a linguistically interpretable
feature of entrenchment and establishes a reproducible procedure for its quantitative assessment based on corpus
observations, without recourse to lexicographic sources.
Янко Т. Е.
Дискурсивное слово ага как маркер пропозициональной установки: функции и просодия
The paper deals with constructions “propositional attitude plus propositional content” including the Russian
discourse word aga ‘aha’. It is shown that aga can either be part of the propositional content of the construction or
function as a propositional attitude as well, representing a predicate of speech or a predicate of a thought derived from
logical inference. These constructions use autonomous prosodic marking of the propositional attitude and
propositional content when the predicate of the propositional attitude or aga used as such bear a fall of fundamental
frequency F0, occasionally combined with emphasis. A communicative-prosodic structure, in which the propositional
attitude and propositional content form an integral speech act, is also possible, but communicative-prosodic autonomy
of propositional attitude and propositional content prevails. To describe the promotion of aga into the position of a
propositional attitude, the concept of functional-semantic insubordination is introduced. The study of aga uncovered
a new prosodic construction not previously recorded in descriptions of Russian prosody. This construction denotes an
extraordinary state of affairs and is characterized by an extremely rapid increase in frequency per unit of time. I refer
to this construction as a “steep rise”.
Зализняк Анна А.
Типы «отношений смежности» в Базе данных семантических переходов в языках мира
The article presents a classification of types of “contiguity relations” reflecting systemic connections
between units of the Database of Semantic Shifts in the Languages of the World. The types of “contiguity
relations” are based on combinations of the following parameters characterizing the relations between the
meanings included in the compared transitions (as a Source and a Target meaning): identity; semantic
proximity; natural binary opposition. We propose to distinguish six types of “related” pairs of shifts and four
types of relations between three shifts. Marking the semantic shifts in the database according to the types of
“contiguity relations” allows us to identify the mechanisms of linguistic conceptualization that are most
consistently implemented in the languages of the world.
Циммерлинг А. В.
Русское отрицание: правила линеаризации и акцентуации
I revise the linearization and accentuation rules for the Russian negative particle ne2 versus the negative
morphemes that can be analyzed as it homographs and follow the distribution of the negative sentence patterns with
sentential and constituent negation in Paducheva’s taxonomy of four negation modes. Both syntactic types of negation
can be interpreted as general or partial negation depending on their contribution to semantic structure. The canonic
types are General Predicate Negation, which brings about verification/falsification semantics and licenses the verum
focus accent on the negative particle, and Constituent Partial Negations which brings about contrastive focusing of a
constituent. The AUX NEG X order can also be interpreted as an instance of General Constituent Negation in noncontrastive and non-verificational contexts. The distribution of four negation modes in Paducheva’s taxonomy is
triggered by the realization of communicative meanings, notably, verification and contrast.

