Open datasets for Kazakh

- NLP and LLM: QA, RAG, instruction tuning, NER, machine translation, and other tasks
- speech: ASR, TTS, speech translation, emotion recognition
- computer vision, OCR, and multimodal tasks
For each dataset, the collection lists its size, scale, task, release year, and download link. It also separately marks resources that were announced but could not yet be independently verified.
Repositories like this make life much easier: instead of spending hours searching, you can quickly see what Kazakh-language data already exists and what you can use for your research or product.
If you know of a dataset that is not yet on the list, the repository is open for contributions: 🔗 https://github.com/Allessyer/awesome-kaz-datasets
Thanks to Asel for this work, and to Alen Issayev for helping develop it!
Comments
Member discussion for this news item or vacancy.