LibreTimes

Works added this month

260 results

Алгебра 1. Лекция 4. Формальные степенные ряды и производящие функции

2026LectureAlexey Ilyin

В этой лекции рассматриваются формальные степенные ряды как алгебраический объект. Вводятся операции сложения, умножения, а также аналоги производной и интеграла. Ключевая идея лекции — применение этого аппарата для решения линейно-рекуррентных соотношений с помощью производящих функций, что демонстрируется на классическом примере чисел Фибоначчи.
0
00

Matrices

2026TheorySergey

A matrix, its notation, and the main types: row, column, echelon, square, diagonal, identity, symmetric - defined and shown with worked examples.
0
00

Reinforcement Networks: novel framework for collaborative Multi-Agent Reinforcement Learning tasks

2025PreprintKryzhanovskiy, Maksim, Glazyrina, Svetlana, Ischenko, Roman, Константин Вячеславович Воронцов

arXiv (Cornell University)

Modern AI systems often comprise multiple learnable components that can be naturally organized as graphs. A central challenge is the end-to-end training of such systems without restrictive architectural or training assumptions. Such tasks fit the theory and approaches of the collaborative Multi-Agent Reinforcement Learning (MARL) field. We introduce Reinforcement Networks, a general framework for MARL that organizes agents as vertices in a directed acyclic graph (DAG). This structure extends hierarchical RL to arbitrary DAGs, enabling flexible credit assignment and scalable coordination while avoiding strict topologies, fully centralized training, and other limitations of current approaches. We formalize training and inference methods for the Reinforcement Networks framework and connect it to the LevelEnv concept to support reproducible construction, training, and evaluation. We demonstrate the effectiveness of our approach on several collaborative MARL setups by developing several Reinforcement Networks models that achieve improved performance over standard MARL baselines. Beyond empirical gains, Reinforcement Networks unify hierarchical, modular, and graph-structured views of MARL, opening a principled path toward designing and training complex multi-agent systems. We conclude with theoretical and practical directions - richer graph morphologies, compositional curricula, and graph-aware exploration. That positions Reinforcement Networks as a foundation for a new line of research in scalable, structured MARL.
0
2

Text Tree Edit Distance: A Language Model-Based Metric for Text Hierarchies

2025Conference paperFedor Sobolevsky, Константин Вячеславович Воронцов

Text trees as a data structure occur in numerous machine learning tasks like hierarchical summarization and automatic mind map generation. One of the main methods of quality evaluation in these tasks is comparison with reference hierarchies created by experts. The method used so far to compare text hierarchies, as shown in this work, poorly accounts for their structure and text semantics relative to phrasing. To address this issue, we propose a new metric on the set of text trees — text tree edit distance (TTED), based on tree edit distance with semantic distance between texts measured using a large language model. To evaluate how the metric reflects different aspects of text tree difference, we introduce special quality coefficients that reflect the sensitivity of a metric to paraphrasing relative to structural and semantic differences of text trees. Using these coefficients, we conduct extensive testing of the proposed metric and its modifications compared to a baseline used in previous works to compare text hierarchies, which shows that TTED indeed captures significant differences between text trees more accurately than the previously used method. We also provide a practical implementation of TTED for further usage.
0
3

ruSciFact: Open Benchmark for Verifying Scientific Facts in Russian

2025Conference paperA. Vatolin, N. Gerasimenko, N. Loukachevitch, A. Ianina +1

Computational Linguistics and Intellectual Technologies

Against the backdrop of active LLM development, their tendency to hallucinate, as well as the growing volume of texts they generate, the validation of facts has become increasingly important and relevant.We propose ruSciFact 1 , a new benchmark for fact-checking scientific claims in Russian.ruSciFact is structured as a NLI task, the goal is to verify whether a fact is confirmed by a given abstract.To generate facts, we used an 3-step pipeline based on LLaMA-405B, validating the resulting sentences with the help of assessors-terminologists.The ruSciFact dataset consists of 1128 pairs in the format , which we are releasing as open source together with the benchmark code.Additionally, we are opensourcing the fact-generation pipeline 2 , which facilitates the expansion of the dataset to specific scientific domains.We evaluated several popular language models on ruSciFact, including text embedders and generative models.The results show that this benchmark allows to effectively assess the fact-checking capabilities of LLMs in Russian.
0
1

The methodology of multi-criteria evaluation of text markup models based on inconsistent expert markup

2025Conference paperAlexander Levikin, Ildar Khabutdinov, Andrey Grabovoy, Константин Вячеславович Воронцов

Computational Linguistics and Intellectual Technologies

A wide class of natural language processing tasks is solved using markup.At the moment, the vast majority of models and datasets rely on a simple markup structure containing only fragments and labels.Moreover, simple classification metrics such as F1, Precision, Recall are used to evaluate the model's accuracy.The problem with such metrics is that they do not take into account all aspects of the markup structure and that they are applicable only under the assumption of the existence of an ideal markup.This paper proposes a more general and universal markup structure that allows solving complex problems and builds a methodology for multi-criteria evaluation of text markup models based on inconsistent expert markup.After that, the application of the constructed method is considered to assess the quality of the model obtained within the winning algorithm of the "READ//ABLE" competition, which focused on building an effective essay markup system.The results demonstrate that the new markup structure and evaluation approach provides a more comprehensive and accurate assessment of model performance, addressing the limitations of traditional metrics by accounting for complex markup scenarios and expert inconsistencies.
0
3

Does Annotating Multi-Spans Improve Classification in Considerable Text Collections?

2024Conference paperArchil Maysuradze, Olga Rink, Artem Fedorov, Andrey Tabachenkov +1

The Universal data markup structures empower the annotation of multiple text fragments (multi-spans or multi-fragments furthermore) to analyze large collections of content. Multi-spans have demonstrated helpful in tackling issues related to the programmed discovery of semantic blunders in school essays or human values in social media writings. Labeling multi-fragment information has made it conceivable to form an interdisciplinary classification of human values. This classifier consists of 105 labels grouped into 7 categories, and a corresponding dataset has been created. Subsequent ML experiments have been designed to demonstrate the effectiveness of the multi-spans structure in recovering annotations of human values. The accuracy of the multi-fragment detector is 0.943 for material values and 0.957 for legal awareness (a subject of civic engagement and citizenship).
0
2

Hypotheses re-ranking in translation models using human markup

2024Journal articleКонстантин Вячеславович Воронцов, N. A. Skachkov

Известия Российской академии наук Теория и системы управления

Modern machine translation systems are trained on large volumes of parallel data obtained using heuristic methods of the Internet bypassing. The poor quality of the data leads to systematic translation errors, which can be quite noticeable from the human point of view. To fix such errors a human based models hypotheses re-ranking is introduced in this work. In this paper the use of human markup is shown not only to increase the overall quality of translation, but also to significantly reduce the number of systematic translation errors. In addition, the relative simplicity of human markup and its integration in the model training process opens up new opportunities in the field of domain adaptation of translation models for new domains like online retail.
0
1

Iterative Improvement of an Additively Regularized Topic Model

2024PreprintGorbulev, Alex, Alekseev, Vasiliy, Константин Вячеславович Воронцов

arXiv (Cornell University)

Topic modelling is fundamentally a soft clustering problem (of known objects -- documents, over unknown clusters -- topics). That is, the task is incorrectly posed. In particular, the topic models are unstable and incomplete. All this leads to the fact that the process of finding a good topic model (repeated hyperparameter selection, model training, and topic quality assessment) can be particularly long and labor-intensive. We aim to simplify the process, to make it more deterministic and provable. To this end, we present a method for iterative training of a topic model. The essence of the method is that a series of related topic models are trained so that each subsequent model is at least as good as the previous one, i.e., that it retains all the good topics found earlier. The connection between the models is achieved by additive regularization. The result of this iterative training is the last topic model in the series, which we call the iteratively updated additively regularized topic model (ITAR). Experiments conducted on several collections of natural language texts show that the proposed ITAR model performs better than other popular topic models (LDA, ARTM, BERTopic), its topics are diverse, and its perplexity (ability to "explain" the underlying data) is moderate.
0
2

LomonosovMSU at SemEval-2024 Task 4: Comparing LLMs and embedder models to identifying propaganda techniques in the content of memes in English for subtasks No1, No2a, and No2b

2024Conference paperGleb Skiba, Mikhail Pukemo, Dmitry Melikhov, Константин Вячеславович Воронцов

This paper presents the solution of the LomonosovMSU team for the SemEval-2024 Task 4 "Multilingual Detection of Persuasion Techniques in Memes" competition for the English language task.During the task solving process, generative and BERT-like (training classifiers on top of embedder models) approaches were tested for subtask №1, as well as an BERT-like approach on top of multimodal embedder models for subtasks №2a/№2b.The models were trained using datasets provided by the competition organizers, enriched with filtered datasets from previous SemEval competitions.The following results were achieved: 18th place for subtask №1, 9th place for subtask №2a, and 11th place for subtask №2b.The code for the solutions is available at github 1 .
0
2

SciRus: Tiny and Powerful Multilingual Encoder for Scientific Texts

2024Journal articleN. Gerasimenko, A. Vatolin, A. Ianina, Константин Вячеславович Воронцов

Doklady Mathematics

LLM-based representation learning is widely used to build effective information retrieval systems, including scientific domains. For making science more open and affordable, it is important that these systems support multilingual (and cross-lingual) search and do not require significant computational power. To address this we propose SciRus-tiny, light multilingual encoder trained from scratch on 44 M abstracts (15B tokens) of research papers and then tuned in a contrastive manner using citation data. SciRus-tiny outperforms SciNCL, English-only SOTA-model for scientific texts, on 13/24 tasks, achieving SOTA on 7, from SciRepEval benchmark. Furthermore, SciRus-tiny is much more effective than SciNCL: it is almost 5x smaller (23 M parameters vs. 110 M), having approximately 2x smaller embeddings (312 vs. 768) and 2x bigger context length (1024 vs. 512). In addition to the tiny model, we also propose the SciRus-small (61 M parameters and 768 embeddings size), which is more powerful and can be used for complicated downstream tasks. We further study different ways of contrastive pre-training and demonstrate that almost SOTA results can be achieved without citation information, operating with only title-abstract pairs.
0
2

RuSciBench: Open Benchmark for Russian and English Scientific Document Representations

2024Journal articleA. Vatolin, N. Gerasimenko, A. Ianina, Константин Вячеславович Воронцов

Doklady Mathematics

Sharing scientific knowledge in the community is an important endeavor. However, most papers are written in English, which makes dissemination of knowledge in countries where English is not spoken by the majority of people harder. Nowadays, machine translation and language models may help to solve this problem, but it is still complicated to train and evaluate models in languages other than English with no or little data in the required language. To address this, we propose the first benchmark for evaluating models on scientific texts in Russian. It consists of papers from Russian electronic library of scientific publications. We also present a set of tasks which can be used to fine-tune various models on our data and provide a detailed comparison between state-of-the-art models on our benchmark.
0
3