LibreTimes

Works

45 results

Hypotheses re-ranking in translation models using human markup

2024Journal articleКонстантин Вячеславович Воронцов, N. A. Skachkov

Известия Российской академии наук Теория и системы управления

Modern machine translation systems are trained on large volumes of parallel data obtained using heuristic methods of the Internet bypassing. The poor quality of the data leads to systematic translation errors, which can be quite noticeable from the human point of view. To fix such errors a human based models hypotheses re-ranking is introduced in this work. In this paper the use of human markup is shown not only to increase the overall quality of translation, but also to significantly reduce the number of systematic translation errors. In addition, the relative simplicity of human markup and its integration in the model training process opens up new opportunities in the field of domain adaptation of translation models for new domains like online retail.
0
1

SciRus: Tiny and Powerful Multilingual Encoder for Scientific Texts

2024Journal articleN. Gerasimenko, A. Vatolin, A. Ianina, Константин Вячеславович Воронцов

Doklady Mathematics

LLM-based representation learning is widely used to build effective information retrieval systems, including scientific domains. For making science more open and affordable, it is important that these systems support multilingual (and cross-lingual) search and do not require significant computational power. To address this we propose SciRus-tiny, light multilingual encoder trained from scratch on 44 M abstracts (15B tokens) of research papers and then tuned in a contrastive manner using citation data. SciRus-tiny outperforms SciNCL, English-only SOTA-model for scientific texts, on 13/24 tasks, achieving SOTA on 7, from SciRepEval benchmark. Furthermore, SciRus-tiny is much more effective than SciNCL: it is almost 5x smaller (23 M parameters vs. 110 M), having approximately 2x smaller embeddings (312 vs. 768) and 2x bigger context length (1024 vs. 512). In addition to the tiny model, we also propose the SciRus-small (61 M parameters and 768 embeddings size), which is more powerful and can be used for complicated downstream tasks. We further study different ways of contrastive pre-training and demonstrate that almost SOTA results can be achieved without citation information, operating with only title-abstract pairs.
0
2

RuSciBench: Open Benchmark for Russian and English Scientific Document Representations

2024Journal articleA. Vatolin, N. Gerasimenko, A. Ianina, Константин Вячеславович Воронцов

Doklady Mathematics

Sharing scientific knowledge in the community is an important endeavor. However, most papers are written in English, which makes dissemination of knowledge in countries where English is not spoken by the majority of people harder. Nowadays, machine translation and language models may help to solve this problem, but it is still complicated to train and evaluate models in languages other than English with no or little data in the required language. To address this, we propose the first benchmark for evaluating models on scientific texts in Russian. It consists of papers from Russian electronic library of scientific publications. We also present a set of tasks which can be used to fine-tune various models on our data and provide a detailed comparison between state-of-the-art models on our benchmark.
0
3

Communicative Type “Municipal Employee” in the Media Space: Development of an Automatic Information and Analytical Assessment System

2024Journal articleIrina Karabulatova, Константин Вячеславович Воронцов, Daniil Okolyshev, Ludan Zhang

Vestnik Volgogradskogo gosudarstvennogo universiteta Serija 2 Jazykoznanije

The article examines the issue of representing municipal government in the media space, followed by the proposed solution for automatically identifying signs of destructive and constructive positioning of communicative types of municipal employees in the public information space. The definition of the concept of the communicative type “municipal employee” with verification features is introduced. The results of the analysis of the organization of local self-government on the example of the Moscow region allowed us to conclude that the communicative type “municipal employee” reflects a diversified system of territorial communicative position within the regional government. The information obtained during the analysis of public information space attitudes regarding the activities of municipal employees can be automated with the method of identifying linguistic markers of emotivity to determine the communicative position of territorial authorities. The suggested methodology for effective automation of the studied subject area in the humanities has been verified as possessing a high scientific potential for further research. It is concluded that the development of technology for monitoring and forecasting public threats based on “soft power” methods through automatic and expert work to identify markers of evaluative presentation of communicative types of municipal employees is designed to help regional authorities achieve the desired results in ensuring territorial identity.
0
3

Reranking Hypotheses in Translation Models Using Human Markup

2024Journal articleКонстантин Вячеславович Воронцов, N. A. Skachkov

Journal of Computer and Systems Sciences International

Modern machine translation systems are trained on large volumes of parallel data obtained using heuristic methods of bypassing the Internet. The poor quality of the data leads to systematic translation errors, which can be quite noticeable to humans. To fix such errors, human-based models for reranking hypotheses is introduced in this study. In this paper the use of human markup is shown not only to increase the overall quality of the translation but also to significantly reduce the number of systematic translation errors. In addition, the relative simplicity of human markup and its integration in the model training process opens up new opportunities in the field of domain adaptation of translation models for new domains like online retail.
0
4

Morphisms of Character Varieties

2024Journal articleSean Cotner

International Mathematics Research Notices

A corrigendum to this paper is published on the author's homepage.
0
2

Verification of communicative types in the judicial public space of media discourse in the USA, Kazakhstan and Russia as a psycholinguistic marker of fact-checking

2023Journal articleGulzat T. Kussepova, Irina S. Karabulatova, Karlygash S. Kenzhigozhina, Aleksey O. Bakhus +1

Revista Amazonia Investiga

Modern psycholinguistic research and fact-checking actively explore the space of media discourse. However, the representation of the judicial space in the mass media has not been sufficiently studied due to the peculiarities of communicative behavior in the judicial and legal space of the ethno-socius and the attitude to the judiciary. The authors hypothesize that the differences in public behavior in court and the coverage of the work of courts in the American, Kazakh and Russian media are due to the socio-cultural features of the phenomena of judicial and legal communication in public space under the influence of established traditions in such coordinate systems as “person – judicial system”, “openness – closeness of society”, “unity – disunity of society”, “accessibility – stigmatization”, “court – journalistic investigation”, etc. The results confirm the hypothesis of the authors' team, revealing the difference in the perception of the judicial system in the USA, Kazakhstan and Russia, illustrating the "rejection" of the Soviet and post-Soviet stigmatization of the judicial and legal space by the Kazakh society towards democratic norms. The prospects of the study are related to the subsequent development of an automatic system for evaluating speech behavior strategies in court and their coverage in the media as a category of fact-checking.
0
2

Система активного наведения для передачи ультрастабильных сигналов оптической частоты по воздушному каналу

2023Journal articleAlexey Legoshin, Ксения Лискова, K. S. Kudeyarov, G. A. Vishnyakova +5

Журнал Экспериментальной и Теоретической Физики

Разработана и создана система активного наведения для атмосферного канала передачи ультрастабильных оптических сигналов частоты, позволяющая существенно уменьшить геометрические отклонения передаваемого лазерного луча и обеспечить стабильную передачу в условиях движущегося отражателя, установленного в средней точке линии. Результаты тестирования работы системы подтверждают ее высокую эффективность и потенциал для применения в реальных условиях.
0
2

Active Pointing System for the Transmission of Ultrastable Optical Frequency Signals through an Open-Air Link

2023Journal articleAlexey Legoshin, Ксения Лискова, K. S. Kudeyarov, G. A. Vishnyakova +5

Journal of Experimental and Theoretical Physics

An active pointing system has been developed and created for an atmospheric transfer link for ultrastable optical frequency signals. This system can significantly decrease the deviations of laser beam direction and ensure stable transmission under conditions of a moving reflector installed at the midpoint of the line. The results of testing the system confirm its high efficiency and potential for use under real conditions.
0
3

Optimizing Modality Weights in Topic Models of Transactional Data

2022Journal articleK. Ya. Khrylchenko, Константин Вячеславович Воронцов

Automation and Remote Control

Modern natural language processing models such as transformers operate multimodal data. In the present paper, multimodal data is explored using multimodal topic modeling on transactional data of bank corporate clients. A definition of the importance of modality for the model is proposed on the basis of which improvements are considered for two modeling scenarios: preserving the maximum amount of information by balancing modalities and automatic selection of modality weights to optimize auxiliary criteria based on topic representations of documents. A model is proposed for adding numerical data to topic models in the form of modalities: each topic is assigned a normal distribution with learning parameters. Significant improvements are demonstrated in comparison with standard topic models on the problem of modeling bank corporate clients. Based on the topic representations of the bank’s customers, a 90-day delay on the loan is predicted.
0
1

Incremental Learning of Topic Models for Finding Trend Topics in Scientific Publications

2022Journal articleN. A. Gerasimenko, A. S. Chernyavsky, M. A. Nikiforova, M. D. Nikitin +1

Doklady Mathematics

With a soaring number of scientific publications and rapid emergence of new directions and approaches, the scientific community faces the task of timely identification of trends. By a trend, we mean a semantically homogeneous topic characterized by a steady lexical kernel and a sharp, often exponential increase in the number of publications [1]. Examples of trends in machine learning are “LSTM,” “deep learning,” “word2vec,” “BERT,” and “fake news detection.” For real-time detection of trend topics from a stream of scientific publications, we use incremental methods of probabilistic topic modeling. An ARTM-based approach to early trend detection has been shown to outperform popular classical and neural network approaches to this task. A dataset of 91 trends for performance evaluation has been manually collected and made available for public use.

0
3

Improving the Quality of Machine Translation Using the Reverse Model

2022Journal articleN. A. Skachkov, Константин Вячеславович Воронцов

Automation and Remote Control

Machine translation is a natural language text processing task that aims to automatically translate input text from one language into another language. The currently known machine translation models show a fairly high quality of translation between large languages, but for smaller language areas, represented by less data, the problem is still not solved. Different methods are used to deal with various errors in automatic translation systems. This paper discusses approaches that use translation models of reverse language directions and improve consistency between translations of the same text using direct and reverse translation models. The paper presents a general theoretical justification for such methods in terms of solving the likelihood maximization problem and also proposes a method for stable training of modern models using cyclic translations.
0
2