LibreTimes

Works

290 results

Convergence of the Algorithm of Additive Regularization of Topic Models

2021Journal articleI. A. Irkhin, Константин Вячеславович Воронцов

Proceedings of the Steklov Institute of Mathematics

The problem of probabilistic topic modeling is as follows. Given a collection of text documents, find the conditional distribution over topics for each document and the conditional distribution over words (or terms) for each topic. Log-likelihood maximization is used to solve this problem. The problem generally has an infinite set of solutions and is ill-posed according to Hadamard. In the framework of Additive Regularization of Topic Models (ARTM), a weighted sum of regularization criteria is added to the main log-likelihood criterion. The numerical method for solving this optimization problem is a kind of an iterative EM-algorithm written in a general form for an arbitrary smooth regularizer as well as for a linear combination of smooth regularizers. This paper studies the problem of convergence of the EM iterative process. Sufficient conditions are obtained for the convergence to a stationary point of the regularized log-likelihood. The constraints imposed on the regularizer are not too restrictive. We give their interpretations from the point of view of the practical implementation of the algorithm. A modification of the algorithm is proposed that improves the convergence without additional time and memory costs. Experiments on a news text collection have shown that our modification both accelerates the convergence and improves the value of the criterion to be optimized.
0
1

Sharpness Estimation of Combinatorial Generalization Ability Bounds for Threshold Decision Rules

2021Journal articleSh. Kh. Ishkina, Константин Вячеславович Воронцов

Automation and Remote Control

This article is devoted to the problem of calculating an exact upper bound for the functionals of the generalization ability of a family of one-dimensional threshold decision rules. An algorithm is investigated that solves the stated problem and is polynomial in the total number of samples used for training and validation and in the number of training samples. A theorem is proved for calculating an estimate for the functional of expected overfitting and an estimate for the error rate of the method for minimizing empirical risk on a validation set. The exact bounds calculated using the theorem are compared with the previously known quick-to-compute upper bounds so as to estimate the orders of overestimation of the bounds and to identify the bounds that could be used in real problems.
0
3

QUANTILE-BASED APPROACH TO ESTIMATING COGNITIVE TEXT COMPLEXITY

2020Conference paperM. A. Eremeev, Константин Вячеславович Воронцов

Computational Linguistics and Intellectual Technologies

This paper introduces an approach to measuring the cognitive complexity of texts on various language levels. While standard readability indices are based on the linear combination of primary statistics, our general approach allows us to estimate complexity on morphological, lexical, syntactic, and discursive levels. Each model is defined by the tokens for the specific language level and the complexity function of a single token. We then use the reference collection of moderately complex texts and the quantile-based approach to spot the abnormally rare tokens. The proposed supervised ensemble, based on the ElasticNet model, incorporates models from all language levels. Having collected a labeled dataset through crowdsourcing, consisting of pairs of articles from the Russian Wikipedia, we consider several models and ensembles and compare them to common baselines. Suggested models are flexible due to the freedom in choosing the reference collection. The described experiments confirm the competitiveness of the proposed approach, as the ensembles demonstrate the best target metric value.
0
4

COMBINING FACTS, SEMANTIC ROLES AND SENTIMENT LEXICON IN A GENERATIVE MODEL FOR OPINION MINING

2020Conference paperD. G. Feldman, T. R. Sadekova, Константин Вячеславович Воронцов

Computational Linguistics and Intellectual Technologies

Opinion mining is a popular task, that is applied, for example, to determine news polarisation and identify product review classes. Our task is unsupervised clusterization of opinionated texts, in particular news on political events. Many papers that tackle this issue use generative models based on lexical features. Our goal is to determine the entities defying an opinion amongst lexical, syntactic and semantic features as well as their compositions. More specifically, we test the hypothesis that an opinion is determined by the composition of the mentioned facts (SPO triples), the semantic roles of the words and the sentiment lexicon used in it. In this paper we formalise this task and prove that using a composition of the above features provides the best quality when clusterising opinionated texts. To test this hypothesis we have gathered and labelled two corpuses of news on political events and proposed a set of unsupervised algorithms for extracting the features.
0
1

Topic Balancing with Additive Regularization of Topic Models

2020Conference paperEugeniia Veselova, Константин Вячеславович Воронцов

This article proposes a new approach for building topic models on unbalanced collections in topic modelling, based on the existing methods and our experiments with such methods. Real-world data collections contain topics in various proportions, and often documents of the relatively small theme become distributed all over the larger topics instead of being grouped into one topic. To address this issue, we design a new regularizer for and matrices in probabilistic Latent Semantic Analysis (pLSA) model. We make sure this regularizer increases the quality of topic models, trained on unbalanced collections. Besides, we conceptually support this regularizer by our experiments.
0
3

Hierarchical Interpretable Topical Embeddings for Exploratory Search and Real-Time Document Tracking

2020Journal articleAnastasia Ianina, Константин Вячеславович Воронцов

International Journal of Embedded and Real-Time Communication Systems

Real-time monitoring of scientific papers and technological news requires fast processing of complicated search demands motivated by thematically relevant information acquisition. For this case, the authors develop an exploratory search engine based on probabilistic hierarchical topic modeling. Topic model gives a low dimensional sparse interpretable vector representation (topical embedding) of a text, which is used for ranking documents by their similarity to the query. They explore several ways of comparing topical vectors including searching with thematically homogeneous text segments. Topical hierarchies are built using the regularized EM-algorithm from BigARTM project. The topic-based search achieves better precision and recall than other approaches (TF-IDF, fastText, LSTM, BERT) and even human assessors who spend up to an hour to complete the same search task. They also discover that blending hierarchical topic vectors with neural pretrained embeddings is a promising way of enriching both models that helps to get precision and recall higher than 90
0
1

Learning Topic Models with Arbitrary Loss

2020Conference paperMurat Apishev, Константин Вячеславович Воронцов

Topic modeling is an area of text analysis actively developing over the past 20 years. A probabilistic topic model (PTM) finds a set of hidden topics from a collection of text documents. It defines each topic as a probability distribution over words and describes each document as a probability mixture of topic distributions. Learning algorithms for topic models are usually based on Bayesian inference or log-likelihood maximization. In both cases, EM-like algorithms are used. In this paper, we propose to replace the logarithm in the log-likelihood by an arbitrary smooth loss function. We prove that such a modification preserves both the structure of the algorithm and compatibility with any regularizers in terms of additive regularization of topic models (ARTM). Moreover, in the case of a linear loss, the Estep becomes much faster due to the omission of a normalization. We study combinations of the fast and usual E-steps and compare them to regularization using different number of topics in both offline and online versions of EM-algorithm. For an empirical comparison of the algorithms, we estimate perplexity, coherence, and learning time. We use an efficient parallel implementation of the EM-algorithm from the BigARTM open-source library. We show that in most cases the two-stage strategy wins, which uses fast E-steps at the beginning of iterations, then proceeds with usual E-steps.
0
4

Движение активной броуновской частицы в сверхтекучем гелии

2020Conference talkAlexey Legoshin, R. E. Boltnev, M. M. Vasiliev, O. F. Petrov

Труды 63-й Всероссийской научной конференции МФТИ 23–29 ноября 2020 года. Фундаментальная и прикладная физика.

В отличие от броуновской частицы в классической жидкости движение такой частицы в сверхтекучем гелии существенно зависит от наличия квантовых вихрей. Известно, что когерентное вращение сверхтекучей компоненты вокруг кора вихря приводит к эффективному захвату примесных частиц вихрями. В этом случае частица либо движется исключительно вдоль кора вихря, либо, если возмущения жидкости достаточно велики, только часть времени проводит в свободном движении между захватами в вихри.

Ситуация качественно изменяется если частица в сверхтекучем гелии оказывается активной, т.е. способной поглощать энергию извне. При достаточно интенсивном тепловыделении частицы у её поверхности формируется противоток нормальной (вязкой) и сверхтекучей компонент, в котором формируются вихри, плотность которых определяется взаимной скоростью нормальной и сверхтекучей компонент, т.е. интенсивностью тепловыделения. Такая частица способна взаимодействовать уже и со сверхтекучей компонентой. Данная работа посвящена исследованию движения активных броуновских частиц в трёхмерном пространстве. В качестве активных частиц были использованы сверхпроводящие частицы с характерным размером 40 мкм, левитирующие в поле магнитной ловушки и облучаемые интенсивным лазерным излучением (~ Вт/см2 ).

0
3

Three-stage question answering system with sentence ranking

2019Conference paperDaria Soboleva, Константин Вячеславович Воронцов

EPiC series in language and linguistics

We explore a recently proposed question answering system. We developed a high speed modification based on dividing the question answering system into three consecutive stages. The first step is to find the candidate documents that most likely contain the answer to the question. The second step is to rank sentences by the probability of having a correct answer to the question. The third step is to find the exact phrase that answers the question. At the third step we used a recently proposed recurrent bidirectional neural network predicting the beginning and the end of a response. In this paper we showed that the proposed question answering system allows to speed up its work without significant losses in the quality. For each step we also explored the feature space construction techniques allowing to improve the final quality.
0
1

Topic Modelling for Extracting Behavioral Patterns from Transactions Data

2019Conference paperEvgeny Egorov, Filipp Nikitin, Vasiliy Alekseev, Alexey Goncharov +1

With the increasing popularity of cashless payment methods for everyday, seasonal and special expenses popular banks accumulate huge amount of data about customer operations. In the article, we report a successful application of topic modelling to extract behaviour patterns from the data. The resulting models are built with BigARTM framework: flexible and efficient tool for topic modelling. The framework allows us to experiment with various models including PLSA, LDA and beyond. Results demonstrate ability of the approach to aggregate information about behaviour patterns of different customer groups. The results analysis allows to see the topics of such people clusters varying from travellers to mortgage holders. Moreover, low-dementional embeddings of the customers, which was given with topic model, were studied. We display that the client vector representations store demographic information as well as source data. We also test for a best way of preparing data for the model with metric above in mind.
0
3

Regularized Multimodal Hierarchical Topic Model for Document-by-Document Exploratory Search

2019Conference paperAnastasia Ianina, Константин Вячеславович Воронцов

In the exploratory search paradigm of information retrieval, the user has a complicated search demand that can not be formulated in a short query. The user collects thematically relevant information iteratively in a “query-browse-refine” process being motivated by learning, understanding, and knowledge acquisition purposes. We consider an elementary step of this scenario in which the search intent can be expressed by a long text query. For this case, we develop an exploratory search engine based on probabilistic topic modeling. Topic model gives a low-dimensional sparse interpretable vector representation (topical embedding) of a text. The search engine uses these embeddings for ranking documents by their similarity to the query. We show that performing only one query, the topic-based search engine achieves better precision and recall that human assessors do spending up to one hour in a conventional browse-refine loop. We use additive regularization for topic modeling (ARTM) to make the model simultaneously sparse, decorrelated, n-gram, multimodal and hierarchical. We show experimentally that each of these features of the model is important to achieve precision and recall higher than 90
0
2

Lexical Quantile-Based Text Complexity Measure

2019Conference paperAITHEA, Russia, Maksim Eremeev, Константин Вячеславович Воронцов

This paper introduces a new approach to estimating the text document complexity.Common readability indices are based on average length of sentences and words.In contrast to these methods, we propose to count the number of rare words occurring abnormally often in the document.We use the reference corpus of texts and the quantile approach in order to determine what words are rare, and what frequencies are abnormal.We construct a general text complexity model, which can be adjusted for the specific task, and introduce two special models.The experimental design is based on a set of thematically similar pairs of Wikipedia articles, labeled using crowdsourcing.The experiments demonstrate the competitiveness of the proposed approach.
0
4

Interpretable probabilistic embeddings: bridging the gap between topic\n models and neural networks

2017PreprintPotapenko, Anna, Popov, Artem, Константин Вячеславович Воронцов

arXiv (Cornell University)

We consider probabilistic topic models and more recent word embedding\ntechniques from a perspective of learning hidden semantic representations.\nInspired by a striking similarity of the two approaches, we merge them and\nlearn probabilistic embeddings with online EM-algorithm on word co-occurrence\ndata. The resulting embeddings perform on par with Skip-Gram Negative Sampling\n(SGNS) on word similarity tasks and benefit in the interpretability of the\ncomponents. Next, we learn probabilistic document embeddings that outperform\nparagraph2vec on a document similarity task and require less memory and time\nfor training. Finally, we employ multimodal Additive Regularization of Topic\nModels (ARTM) to obtain a high sparsity and learn embeddings for other\nmodalities, such as timestamps and categories. We observe further improvement\nof word similarity performance and meaningful inter-modality similarities.\n

0
4