LibreTimes

Search results for “computer-science”

30 results

Reranking Hypotheses in Translation Models Using Human Markup

2024Journal articleK. V. Vorontsov, N. A. Skachkov

Journal of Computer and Systems Sciences International

Modern machine translation systems are trained on large volumes of parallel data obtained using heuristic methods of bypassing the Internet. The poor quality of the data leads to systematic translation errors, which can be quite noticeable to humans. To fix such errors, human-based models for reranking hypotheses is introduced in this study. In this paper the use of human markup is shown not only to increase the overall quality of the translation but also to significantly reduce the number of systematic translation errors. In addition, the relative simplicity of human markup and its integration in the model training process opens up new opportunities in the field of domain adaptation of translation models for new domains like online retail.
0
1

Regularization, robustness and sparsity of probabilistic topic models

2012Journal articleKonstantin Vyacheslavovich Vorontsov, Anna Alexandrovna Potapenko

Computer Research and Modeling

We propose a generalized probabilistic topic model of text corpora which can incorporate heuristics of Bayesian regularization, sampling, frequent parameters update, and robustness in any combinations. Wellknown models PLSA, LDA, CVB0, SWB, and many others can be considered as special cases of the proposed broad family of models. We propose the robust PLSA model and show that it is more sparse and performs better that regularized models like LDA.
0
1