Search results for “sine”
Categories
170 results
QUANTILE-BASED APPROACH TO ESTIMATING COGNITIVE TEXT COMPLEXITY
2020Conference paperM. A. Eremeev, K. V. Vorontsov
Computational Linguistics and Intellectual Technologies
This paper introduces an approach to measuring the cognitive complexity of texts on various language levels. While standard readability indices are based on the linear combination of primary statistics, our general approach allows us to estimate complexity on morphological, lexical, syntactic, and discursive levels. Each model is defined by the tokens for the specific language level and the complexity function of a single token. We then use the reference collection of moderately complex texts and the quantile-based approach to spot the abnormally rare tokens. The proposed supervised ensemble, based on the ElasticNet model, incorporates models from all language levels. Having collected a labeled dataset through crowdsourcing, consisting of pairs of articles from the Russian Wikipedia, we consider several models and ensembles and compare them to common baselines. Suggested models are flexible due to the freedom in choosing the reference collection. The described experiments confirm the competitiveness of the proposed approach, as the ensembles demonstrate the best target metric value.
0
1
SciRus: Tiny and Powerful Multilingual Encoder for Scientific Texts
2024Journal articleN. Gerasimenko, A. Vatolin, A. Ianina, K. Vorontsov
Doklady Mathematics
LLM-based representation learning is widely used to build effective information retrieval systems, including scientific domains. For making science more open and affordable, it is important that these systems support multilingual (and cross-lingual) search and do not require significant computational power. To address this we propose SciRus-tiny, light multilingual encoder trained from scratch on 44 M abstracts (15B tokens) of research papers and then tuned in a contrastive manner using citation data. SciRus-tiny outperforms SciNCL, English-only SOTA-model for scientific texts, on 13/24 tasks, achieving SOTA on 7, from SciRepEval benchmark. Furthermore, SciRus-tiny is much more effective than SciNCL: it is almost 5x smaller (23 M parameters vs. 110 M), having approximately 2x smaller embeddings (312 vs. 768) and 2x bigger context length (1024 vs. 512). In addition to the tiny model, we also propose the SciRus-small (61 M parameters and 768 embeddings size), which is more powerful and can be used for complicated downstream tasks. We further study different ways of contrastive pre-training and demonstrate that almost SOTA results can be achieved without citation information, operating with only title-abstract pairs.
0
1
IMO 1967
Problem sheet
International mathematics competition for pre-university students, held annually since 1959. Six problems over two days.
0
0
IMO 1970
Problem sheet
International mathematics competition for pre-university students, held annually since 1959. Six problems over two days.
0
0