LibreTimes

Search results for “representation theory”

Categories

6 results

SciRus: Tiny and Powerful Multilingual Encoder for Scientific Texts

2024Journal articleN. Gerasimenko, A. Vatolin, A. Ianina, K. Vorontsov

Doklady Mathematics

LLM-based representation learning is widely used to build effective information retrieval systems, including scientific domains. For making science more open and affordable, it is important that these systems support multilingual (and cross-lingual) search and do not require significant computational power. To address this we propose SciRus-tiny, light multilingual encoder trained from scratch on 44 M abstracts (15B tokens) of research papers and then tuned in a contrastive manner using citation data. SciRus-tiny outperforms SciNCL, English-only SOTA-model for scientific texts, on 13/24 tasks, achieving SOTA on 7, from SciRepEval benchmark. Furthermore, SciRus-tiny is much more effective than SciNCL: it is almost 5x smaller (23 M parameters vs. 110 M), having approximately 2x smaller embeddings (312 vs. 768) and 2x bigger context length (1024 vs. 512). In addition to the tiny model, we also propose the SciRus-small (61 M parameters and 768 embeddings size), which is more powerful and can be used for complicated downstream tasks. We further study different ways of contrastive pre-training and demonstrate that almost SOTA results can be achieved without citation information, operating with only title-abstract pairs.
0
1

Verification of communicative types in the judicial public space of media discourse in the USA, Kazakhstan and Russia as a psycholinguistic marker of fact-checking

2023Journal articleGulzat T. Kussepova, Irina S. Karabulatova, Karlygash S. Kenzhigozhina, Aleksey O. Bakhus +1

Revista Amazonia Investiga

Modern psycholinguistic research and fact-checking actively explore the space of media discourse. However, the representation of the judicial space in the mass media has not been sufficiently studied due to the peculiarities of communicative behavior in the judicial and legal space of the ethno-socius and the attitude to the judiciary. The authors hypothesize that the differences in public behavior in court and the coverage of the work of courts in the American, Kazakh and Russian media are due to the socio-cultural features of the phenomena of judicial and legal communication in public space under the influence of established traditions in such coordinate systems as “person – judicial system”, “openness – closeness of society”, “unity – disunity of society”, “accessibility – stigmatization”, “court – journalistic investigation”, etc. The results confirm the hypothesis of the authors' team, revealing the difference in the perception of the judicial system in the USA, Kazakhstan and Russia, illustrating the "rejection" of the Soviet and post-Soviet stigmatization of the judicial and legal space by the Kazakh society towards democratic norms. The prospects of the study are related to the subsequent development of an automatic system for evaluating speech behavior strategies in court and their coverage in the media as a category of fact-checking.
0
1

Hierarchical Interpretable Topical Embeddings for Exploratory Search and Real-Time Document Tracking

2020Journal articleAnastasia Ianina, Konstantin Vorontsov

International Journal of Embedded and Real-Time Communication Systems

Real-time monitoring of scientific papers and technological news requires fast processing of complicated search demands motivated by thematically relevant information acquisition. For this case, the authors develop an exploratory search engine based on probabilistic hierarchical topic modeling. Topic model gives a low dimensional sparse interpretable vector representation (topical embedding) of a text, which is used for ranking documents by their similarity to the query. They explore several ways of comparing topical vectors including searching with thematically homogeneous text segments. Topical hierarchies are built using the regularized EM-algorithm from BigARTM project. The topic-based search achieves better precision and recall than other approaches (TF-IDF, fastText, LSTM, BERT) and even human assessors who spend up to an hour to complete the same search task. They also discover that blending hierarchical topic vectors with neural pretrained embeddings is a promising way of enriching both models that helps to get precision and recall higher than 90
0
1

RuSciBench: Open Benchmark for Russian and English Scientific Document Representations

2024Journal articleA. Vatolin, N. Gerasimenko, A. Ianina, K. Vorontsov

Doklady Mathematics

Sharing scientific knowledge in the community is an important endeavor. However, most papers are written in English, which makes dissemination of knowledge in countries where English is not spoken by the majority of people harder. Nowadays, machine translation and language models may help to solve this problem, but it is still complicated to train and evaluate models in languages other than English with no or little data in the required language. To address this, we propose the first benchmark for evaluating models on scientific texts in Russian. It consists of papers from Russian electronic library of scientific publications. We also present a set of tasks which can be used to fine-tune various models on our data and provide a detailed comparison between state-of-the-art models on our benchmark.
0
1

Optimizing Modality Weights in Topic Models of Transactional Data

2022Journal articleK. Ya. Khrylchenko, K. V. Vorontsov

Automation and Remote Control

Modern natural language processing models such as transformers operate multimodal data. In the present paper, multimodal data is explored using multimodal topic modeling on transactional data of bank corporate clients. A definition of the importance of modality for the model is proposed on the basis of which improvements are considered for two modeling scenarios: preserving the maximum amount of information by balancing modalities and automatic selection of modality weights to optimize auxiliary criteria based on topic representations of documents. A model is proposed for adding numerical data to topic models in the form of modalities: each topic is assigned a normal distribution with learning parameters. Significant improvements are demonstrated in comparison with standard topic models on the problem of modeling bank corporate clients. Based on the topic representations of the bank’s customers, a 90-day delay on the loan is predicted.
0
1