LibreTimes

Search results for “representation theory”

Categories

10 results

The Steinberg Representation

2026TheorySean Cotner

An introduction to the Steinberg representation of a finite group of Lie type — its alternating-sum construction from parabolic inductions, worked out explicitly for SL_2.

0
00

SciRus: Tiny and Powerful Multilingual Encoder for Scientific Texts

2024Journal articleN. Gerasimenko, A. Vatolin, A. Ianina, K. Vorontsov

Doklady Mathematics

LLM-based representation learning is widely used to build effective information retrieval systems, including scientific domains. For making science more open and affordable, it is important that these systems support multilingual (and cross-lingual) search and do not require significant computational power. To address this we propose SciRus-tiny, light multilingual encoder trained from scratch on 44 M abstracts (15B tokens) of research papers and then tuned in a contrastive manner using citation data. SciRus-tiny outperforms SciNCL, English-only SOTA-model for scientific texts, on 13/24 tasks, achieving SOTA on 7, from SciRepEval benchmark. Furthermore, SciRus-tiny is much more effective than SciNCL: it is almost 5x smaller (23 M parameters vs. 110 M), having approximately 2x smaller embeddings (312 vs. 768) and 2x bigger context length (1024 vs. 512). In addition to the tiny model, we also propose the SciRus-small (61 M parameters and 768 embeddings size), which is more powerful and can be used for complicated downstream tasks. We further study different ways of contrastive pre-training and demonstrate that almost SOTA results can be achieved without citation information, operating with only title-abstract pairs.
0
1

Verification of communicative types in the judicial public space of media discourse in the USA, Kazakhstan and Russia as a psycholinguistic marker of fact-checking

2023Journal articleGulzat T. Kussepova, Irina S. Karabulatova, Karlygash S. Kenzhigozhina, Aleksey O. Bakhus +1

Revista Amazonia Investiga

Modern psycholinguistic research and fact-checking actively explore the space of media discourse. However, the representation of the judicial space in the mass media has not been sufficiently studied due to the peculiarities of communicative behavior in the judicial and legal space of the ethno-socius and the attitude to the judiciary. The authors hypothesize that the differences in public behavior in court and the coverage of the work of courts in the American, Kazakh and Russian media are due to the socio-cultural features of the phenomena of judicial and legal communication in public space under the influence of established traditions in such coordinate systems as “person – judicial system”, “openness – closeness of society”, “unity – disunity of society”, “accessibility – stigmatization”, “court – journalistic investigation”, etc. The results confirm the hypothesis of the authors' team, revealing the difference in the perception of the judicial system in the USA, Kazakhstan and Russia, illustrating the "rejection" of the Soviet and post-Soviet stigmatization of the judicial and legal space by the Kazakh society towards democratic norms. The prospects of the study are related to the subsequent development of an automatic system for evaluating speech behavior strategies in court and their coverage in the media as a category of fact-checking.
0
1

Regularized Multimodal Hierarchical Topic Model for Document-by-Document Exploratory Search

2019Conference paperAnastasia Ianina, Konstantin Vorontsov

In the exploratory search paradigm of information retrieval, the user has a complicated search demand that can not be formulated in a short query. The user collects thematically relevant information iteratively in a “query-browse-refine” process being motivated by learning, understanding, and knowledge acquisition purposes. We consider an elementary step of this scenario in which the search intent can be expressed by a long text query. For this case, we develop an exploratory search engine based on probabilistic topic modeling. Topic model gives a low-dimensional sparse interpretable vector representation (topical embedding) of a text. The search engine uses these embeddings for ranking documents by their similarity to the query. We show that performing only one query, the topic-based search engine achieves better precision and recall that human assessors do spending up to one hour in a conventional browse-refine loop. We use additive regularization for topic modeling (ARTM) to make the model simultaneously sparse, decorrelated, n-gram, multimodal and hierarchical. We show experimentally that each of these features of the model is important to achieve precision and recall higher than 90
0
1

Hierarchical Interpretable Topical Embeddings for Exploratory Search and Real-Time Document Tracking

2020Journal articleAnastasia Ianina, Konstantin Vorontsov

International Journal of Embedded and Real-Time Communication Systems

Real-time monitoring of scientific papers and technological news requires fast processing of complicated search demands motivated by thematically relevant information acquisition. For this case, the authors develop an exploratory search engine based on probabilistic hierarchical topic modeling. Topic model gives a low dimensional sparse interpretable vector representation (topical embedding) of a text, which is used for ranking documents by their similarity to the query. They explore several ways of comparing topical vectors including searching with thematically homogeneous text segments. Topical hierarchies are built using the regularized EM-algorithm from BigARTM project. The topic-based search achieves better precision and recall than other approaches (TF-IDF, fastText, LSTM, BERT) and even human assessors who spend up to an hour to complete the same search task. They also discover that blending hierarchical topic vectors with neural pretrained embeddings is a promising way of enriching both models that helps to get precision and recall higher than 90
0
1

RuSciBench: Open Benchmark for Russian and English Scientific Document Representations

2024Journal articleA. Vatolin, N. Gerasimenko, A. Ianina, K. Vorontsov

Doklady Mathematics

Sharing scientific knowledge in the community is an important endeavor. However, most papers are written in English, which makes dissemination of knowledge in countries where English is not spoken by the majority of people harder. Nowadays, machine translation and language models may help to solve this problem, but it is still complicated to train and evaluate models in languages other than English with no or little data in the required language. To address this, we propose the first benchmark for evaluating models on scientific texts in Russian. It consists of papers from Russian electronic library of scientific publications. We also present a set of tasks which can be used to fine-tune various models on our data and provide a detailed comparison between state-of-the-art models on our benchmark.
0
1

Interpretable probabilistic embeddings: bridging the gap between topic\n models and neural networks

2017PreprintPotapenko, Anna, Popov, Artem, Vorontsov, Konstantin

arXiv (Cornell University)

We consider probabilistic topic models and more recent word embedding\ntechniques from a perspective of learning hidden semantic representations.\nInspired by a striking similarity of the two approaches, we merge them and\nlearn probabilistic embeddings with online EM-algorithm on word co-occurrence\ndata. The resulting embeddings perform on par with Skip-Gram Negative Sampling\n(SGNS) on word similarity tasks and benefit in the interpretability of the\ncomponents. Next, we learn probabilistic document embeddings that outperform\nparagraph2vec on a document similarity task and require less memory and time\nfor training. Finally, we employ multimodal Additive Regularization of Topic\nModels (ARTM) to obtain a high sparsity and learn embeddings for other\nmodalities, such as timestamps and categories. We observe further improvement\nof word similarity performance and meaningful inter-modality similarities.\n

0
1

Topic Modelling for Extracting Behavioral Patterns from Transactions Data

2019Conference paperEvgeny Egorov, Filipp Nikitin, Vasiliy Alekseev, Alexey Goncharov +1

With the increasing popularity of cashless payment methods for everyday, seasonal and special expenses popular banks accumulate huge amount of data about customer operations. In the article, we report a successful application of topic modelling to extract behaviour patterns from the data. The resulting models are built with BigARTM framework: flexible and efficient tool for topic modelling. The framework allows us to experiment with various models including PLSA, LDA and beyond. Results demonstrate ability of the approach to aggregate information about behaviour patterns of different customer groups. The results analysis allows to see the topics of such people clusters varying from travellers to mortgage holders. Moreover, low-dementional embeddings of the customers, which was given with topic model, were studied. We display that the client vector representations store demographic information as well as source data. We also test for a best way of preparing data for the model with metric above in mind.
0
1

Incremental Topic Modeling for Scientific Trend Topics Extraction

2023Conference paperNikolai Gerasimenko, Alexander Chernyavskiy, Maria Nikiforova, Anastasia Ianina +1

Computational Linguistics and Intellectual Technologies

Rapid growth of scientific publications and intensive emergence of new directions and approaches poses a challenge to the scientific community to identify trends in a timely and automatic manner. We denote trend as a semantically homogeneous theme that is characterized by a lexical kernel steadily evolving in time and a sharp, often exponential, increase in the number of publications. In this paper, we investigate recent topic modeling approaches to accurately extract trending topics at an early stage. In particular, we customize the standard ARTM-based approach and propose a novel incremental training technique which helps the model to operate on data in real-time. We further create the Artificial Intelligence Trends Dataset (AITD) that contains a collection of early-stage articles and a set of key collocations for each trend. The conducted experiments demonstrate that the suggested ARTM-based approach outperforms the classic PLSA, LDA models and a neural approach based on BERT representations. Our models and dataset are open for research purposes.
0
1

Optimizing Modality Weights in Topic Models of Transactional Data

2022Journal articleK. Ya. Khrylchenko, K. V. Vorontsov

Automation and Remote Control

Modern natural language processing models such as transformers operate multimodal data. In the present paper, multimodal data is explored using multimodal topic modeling on transactional data of bank corporate clients. A definition of the importance of modality for the model is proposed on the basis of which improvements are considered for two modeling scenarios: preserving the maximum amount of information by balancing modalities and automatic selection of modality weights to optimize auxiliary criteria based on topic representations of documents. A model is proposed for adding numerical data to topic models in the form of modalities: each topic is assigned a normal distribution with learning parameters. Significant improvements are demonstrated in comparison with standard topic models on the problem of modeling bank corporate clients. Based on the topic representations of the bank’s customers, a 90-day delay on the loan is predicted.
0
1