TGViewer
adiga.ai adiga.ai @adiga_ai · 409 subscribers
Post #8 10.2K
Завершена работа над первой версией датасета русско-черкесских параллельных текстов. Датасет состоит из около 330 тысяч пар переводов: 220 тысяч на восточном (кабардинском) диалекте и 110 тысяч на западном. Тексты собирались в течение нескольких лет из различных словарей, книг, статей, а также с помощью волонтеров на zedzek.com. Спасибо всем кто принимает участие в сборе данных.

Датасет опубликован в открытом доступе на Hugging Face: https://huggingface.co/datasets/adiga-ai/circassian-parallel-corpus
Любой желающий может использовать его для обучения моделей, в академических и любых других целях.

Главной целью проекта adiga.ai является расширение присутствия черкесского языка в интернете. Поэтому датасет также был передан представителям компаний Яндекс, Гугл и Мета, которые планируют использовать его для обучения своих мультиязычных моделей. Если все пойдет хорошо, то в течение ближайшего года можно рассчитывать на появление черкесского языка в Яндекс Переводчике, Google Переводчике и его поддержку в продуктах компании Meta (facebook, instagram), а также в открытых языковых моделях этих компаний.

* * *

The first version of the Russian-Circassian parallel text dataset has been completed. The dataset consists of ~330,000 translation pairs: 220,000 in the Eastern (Kabardian) dialect and 110,000 in the Western dialect. These texts were compiled over several years from various dictionaries, books, and articles, as well as through contributions from volunteers at zedzek.com. Thanks a lot to everyone who contributed to collecting the data.

The dataset has been made publicly available on Hugging Face:
https://huggingface.co/datasets/adiga-ai/circassian-parallel-corpus
Anyone interested is free to use it for model training, academic research, or any other purposes.

The primary goal of the adiga.ai project is to increase the presence of the Circassian language online. To support this goal, the dataset has also been shared with representatives from Yandex, Google, and Meta, who plan to use it as part of their ongoing projects to train multilingual models. If everything goes well, we can expect Circassian to become available in Yandex Translate, Google Translate, and supported across Meta products (Facebook, Instagram), as well as integrated into open-source language models from these companies within the coming year.
  • 🔥 30
  • ❤ 17
  • 🙏 5
  • ❤‍🔥 3
  • 💘 1
More from @adiga_ai
  1. Sep 25, 2025Прошло чуть больше месяца как инженерам Яндекса был передан датасет adiga.ai из ~300 тыс.…
  2. Sep 25, 2025Just over a month has passed since Yandex engineers received the adiga.ai dataset containi…
  3. May 28, 2025A small quality-of-life update: Both zedzek.com and adiga.ai/chat now have a toggleable vi…
  4. May 1, 2025Черкесский виртуальный помощник Нарт На сайте adiga.ai/chat запущен черкесский чатбот, ана…
  5. May 1, 2025Большое обновление переводчика zedzek.com Значительно улучшено качество перевода: сайт раб…
  6. May 1, 2025adiga.ai – проект, в рамках которого я надеюсь внести вклад в сохранение и популяризацию ч…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →