Latest news and updates from DSML Kazakhstan community
Stay up to date with the latest events
DSML Reading Club is a series of informal online community meetings where members share interesting ideas, recent papers, and their own research. If you recently read a great paper, published your work in a journal or at a conference, or simply want to discuss an interesting topic, join us as a speaker!
Complete the registration form. We will contact you and help organize the meeting.
Beetech 2025 took place last Saturday. Alongside the General track, this year's conference featured an AI & Beyond track, where members of our community delivered half of the talks.
Abylaikhan Turlasov shared guided-decoding techniques for working with LLMs. Mikhail Shkorin explained how to prevent video substitution during authentication. Dias Khalniyazov discussed the details of building an AI parser, while Renat Alimbekov broke down common mistakes in AI/ML projects and answered an important question: Is Lockheed real?
The day after tomorrow, Assel Yermekova will present a paper she co-authored, “Improved Sampling Algorithms for Lévy-Itô Diffusion Models”.
Recent work showed that Lévy–Itô diffusion models with isotropic α-stable noise improve image generation on imbalanced data. However, existing sampling algorithms solve only approximate reverse equations, which reduces quality. In this paper, we propose a family of stochastic differential equations with identical marginal distributions and show that parameter selection improves quality with few reverse-diffusion steps. We also demonstrate Lévy–Itô models across different domains and the advantages of text-to-speech models on highly imbalanced data.
At the meeting, we will discuss:
Anuar Taskynov presented Visual Geometry Grounded Transformer.
VGGT is a next-generation foundation model for 3D computer-vision tasks. From one, several, or even hundreds of scene images, it can immediately predict key 3D properties: camera parameters, depth maps, dense point clouds, and 3D tracking.
Unlike traditional approaches, VGGT works as a single universal model without complex post-processing. It remains fast—under one second per reconstruction—and accurate, achieving state-of-the-art results across several 3D tasks.
Seminar host: Yelaman Abdullin. Download the presentation.
Watch the video: youtube.com/watch?v=TVZoU1m5WKI
Yelaman Abdullin presented Byte Latent Transformer.
Modern LLMs rely on tokenization, which limits flexibility, reduces efficiency, and makes them vulnerable to rare and irregular inputs. The paper proposes Byte Latent Transformer (BLT), a new architecture that works directly with bytes. BLT uses dynamic patches that adapt to data complexity and, for the first time, matches the quality of tokenized models while providing better efficiency and scalability.
Watch the video: youtu.be/JN-adAvbAcs