Project Vakyansh
The first and largest open-source ASR system for Indian languages, adopted by the Indian government. Most speech data and most speech models assume a handful of languages; this was an attempt to widen that.
I build machine learning systems that reach people — speech models in 23 languages, health agents on wearable data, and the evaluation work that decides whether any of it is trustworthy.
At Microsoft I work on AI agents and large language models, and on the evaluation frameworks that tell us whether a production system is actually doing what we think it is. Evaluation is the least glamorous part of applied ML and usually the part that decides whether a model ships.
Before that I spent five years building production systems across speech, health, and mobility — most of it the unglamorous middle ground between a research result and something people can rely on.
I speak on applied LLM systems, evaluation, and building ML for languages and populations that get left out of benchmarks. Open to conference talks, panels, and hackathon judging — the fastest way to reach me is LinkedIn.
Eight or more publications and 90+ citations across speech recognition, NLP, self-supervised learning, and large language models, including work at ICLR 2024.
The first and largest open-source ASR system for Indian languages, adopted by the Indian government. Most speech data and most speech models assume a handful of languages; this was an attempt to widen that.
I write and teach AI on Instagram at @ai.priyanshi, mostly for people trying to break into the field. Professional history and contact live on LinkedIn.