Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Fine-Tune SraVaani on Your Own Speech Data

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Sujith Pulikodan, Agneedh Basu, PavanKumarJ, Pranav D Bhat, and Suryansh Shukla
Word Count
2,759
Company Posts That Month
74
Language
-
Hacker News Points
-
Post removed?
No
Summary

Published by ARTPARK-IISc, the guide explains how to fine-tune SraVaani, a Hybrid RNN-T and CTC FastConformer speech-recognition model pretrained on dozens of Indian languages, for a new language or domain using custom audio and transcripts on a single GPU with at least 15GB of VRAM. Using the low-resource Wancho language as an example, it covers environment setup with the required CUDA-enabled NeMo installation, checkpoint downloading and integrity verification, preparation of 16 kHz mono audio and JSONL manifests, and the choice between ordinary audio files and tarred shards for larger datasets. It outlines loading the model, deciding whether to freeze the encoder for small datasets or fully fine-tune it, configuring hybrid decoder loss, conservative optimization settings, training, checkpointing, monitoring, and evaluating original and adapted models with consistently normalized word error rate scores. In the example, a two-epoch decoder-focused run on several hours of Wancho audio reduced test WER from 65.26% to 64.22%, which the authors present as a modest but useful validation of the pipeline; they suggest that more data, longer training, or carefully unfreezing the encoder may yield larger improvements.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 6 516 143 56 -47%
LLM 1 4,718 960 222 -38%
Voice AI 1 2,814 261 53 -37%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.