Open datasets for Kazakh

Open in Telegram
Open datasets for Kazakh
  • NLP and LLM: QA, RAG, instruction tuning, NER, machine translation, and other tasks
  • speech: ASR, TTS, speech translation, emotion recognition
  • computer vision, OCR, and multimodal tasks

For each dataset, the collection lists its size, scale, task, release year, and download link. It also separately marks resources that were announced but could not yet be independently verified.

Repositories like this make life much easier: instead of spending hours searching, you can quickly see what Kazakh-language data already exists and what you can use for your research or product.

If you know of a dataset that is not yet on the list, the repository is open for contributions: 🔗 https://github.com/Allessyer/awesome-kaz-datasets

Thanks to Asel for this work, and to Alen Issayev for helping develop it!

Comments

Member discussion for this news item or vacancy.

Checking sign-in status...