Breaking News • AI • Technology • Startups • Cybersecurity • Future Tech

Sarvam AI Unveils Saaras V4: A Breakthrough Speech-to-Text Model for All 22 Indian Languages and Global English

Sarvam AI Unveils Saaras V4: A Breakthrough Speech-to-Text Model for All 22 Indian Languages and Global English

Revolutionizing Speech Recognition for India and Beyond

For readers tracking the shift, Sarvam AI has announced the release of Saaras V4, the latest iteration of its advanced speech recognition model. This groundbreaking development marks a significant leap forward, offering state-of-the-art accuracy across all 22 officially recognized Indian languages, alongside comprehensive support for global English accents. Saaras V4 is poised to transform how businesses and developers interact with spoken language, especially in India’s diverse linguistic landscape.

Unmatched Linguistic Coverage

Meanwhile, One of Saaras V4’s most compelling features is its unparalleled language support. For the first time, a single model from Sarvam AI accurately transcribes all 22 scheduled Indian languages, making it an indispensable tool for applications targeting India’s vast and multilingual population. Furthermore, its enhanced capability to process global English accents ensures broad applicability beyond national borders, catering to a wider international audience.

How Saaras V4 Works: An Inside Look

At its core, Saaras V4 operates as an sophisticated encoder-decoder system. Here’s a simplified breakdown:

  • Audio Encoder: This component transforms raw audio waveforms into detailed embeddings, capturing crucial phonetic and acoustic information.
  • Temporal-Downsampling Adapter: To efficiently manage longer recordings, this adapter shortens the sequence of embeddings and projects them into the language model’s embedding space. This ensures that even extended audio fits within the decoder’s processing capacity.
  • Sarvam-3B Decoder: The heart of the system is Sarvam-3B, a powerful 3-billion parameter hybrid state-space language model developed from scratch in-house. It processes the audio features in conjunction with a text prompt, then autoregressively generates the transcript, feeding each token back to predict the next.

Exceptional Performance Across Diverse Benchmarks

In practical terms, Sarvam AI reports Saaras V4 delivers impressive accuracy, outperforming several competitors in various benchmarks. Notably these figures are vendor-reported, and independent verification is pending.

English Language Performance

Evaluated across seven distinct English datasets, including those from Hugging Face’s Open ASR Leaderboard (AMI, GigaSpeech, LibriSpeech clean/other, SPGISpeech, VoxPopuli) and AI4Bharat’s Indian-accented Svarah, Saaras V4 achieved the lowest average Word Error Rate (WER) among the models benchmarked by Sarvam.

Indic Language Accuracy

For example, On the Vistaar dataset, Saaras V4 demonstrated strong results across 10 Indian languages. Sarvam utilized both WER and LLM-WER, with the latter providing a semantic check to differentiate true meaning errors from minor spelling or formatting variations common in Indic scripts.

Robustness in Noisy Environments

For challenging audio, such as compressed, clipped, or background-heavy recordings found in the Kathbath Noisy dataset, Saaras V4 significantly reduced error rates (measured with LLM-WER) to less than half that of leading competitors like Deepgram Nova-3 and GPT-4o Transcribe.

Precise Language Identification

That said, Saaras V4 also excels at identifying languages. On verified IndicVoices utterances, it achieved a language identification error rate of just 2.9% across the top 10 Indian languages, extending to 5.22% across all 22 supported languages.

Five Dynamic Output Modes from a Single Model

Saaras V4 offers unprecedented flexibility with five distinct output modes, all accessible from the same model, eliminating the need for complex post-processing:

  1. Transcribe (default): Provides the native script with normalized numbers and dates.
  2. Verbatim: Delivers every spoken word, including fillers and spoken numbers.
  3. Codemix: Presents the native script, while retaining English words in their original English form.
  4. Translit: Outputs the entire utterance in Latin script.
  5. Translate: Generates an English translation with normalized numbers.

Interestingly, This integrated approach significantly reduces the potential for errors that can arise from external post-processing steps.

Introducing Keyterm Prompting

A significant new feature in Saaras V4 is keyterm prompting. Users can provide a JSON list of up to 50 terms (each up to 64 characters) to bias the recognition process. While not a guarantee, this feature significantly improves the model’s ability to accurately transcribe specific words or phrases.

For instance, using codemix mode alongside keyterm prompting ensures brand names like ‘PhonePe’ remain in Latin script. On the IndicContextEval (L5 keyword-prompting setting), Saaras V4 achieved a reported WER of 16.03%, the lowest on this benchmark.

Flexible Deployment and Accessibility

Saaras V4 is designed for versatile deployment, catering to various application needs:

  • Streaming: Utilizes WebSockets for real-time transcription, boasting a time to first token (TTFT) below 150 ms.
  • REST: Offers synchronous transcription for audio clips up to 30 seconds.
  • Batch: Supports asynchronous jobs for files up to 2 hours, with optional speaker diarization.
  • SDKs: Available for Python 3.9+ and Node.js 18+, with integrations for LiveKit Agents, Pipecat, and Vercel AI SDK.

Sarvam AI has priced speech-to-text services at an affordable ₹30 per hour for real-time, streaming, and batch processing, with diarization available at ₹45 per hour. While Saaras V3 remains the default model, upgrading to V4 requires just a single line of code change, thanks to its identical request shape.

Saaras V4: A Competitive Edge

Meanwhile, Compared to leading competitors like Deepgram Nova-3, ElevenLabs Scribe v2, and OpenAI GPT-4o Transcribe, Saaras V4 stands out with its comprehensive support for all 22 Indian scheduled languages and its unique suite of five built-in output modes. While self-hosting for V4 is not yet available (V3 is on SageMaker), its cloud API offers a compelling and cost-effective solution for multilingual speech-to-text needs.

Key Takeaways for Developers and Businesses

  • Broad Language Support: Covers all 22 Indian languages plus global English in a single, highly accurate model.
  • Advanced Architecture: Powered by an in-house trained 3B hybrid state-space decoder.
  • Enhanced Accuracy: Achieves state-of-the-art results on various benchmarks, including noisy audio and Indic languages (vendor-reported).
  • Flexible Outputs: Offers five distinct output modes, streamlining post-processing.
  • Smart Prompting: New keyterm prompting feature improves recognition of specific terms.
  • Accessible Deployment: Available via API with streaming, REST, and batch options, and competitive pricing.

Saaras V4 represents a significant advancement in speech recognition technology, particularly for the Indian market, offering powerful, flexible, and accurate transcription capabilities.

Expert Perspective

From an industry angle, the clearest signal around Sarvam AI Saaras V4 is how it may influence saaras. The story reads less like a one-day spike and more like a marker of broader movement.

The next phase will depend on how quickly teams, regulators, or customers react. In practice, that gives Sarvam AI Saaras V4 room to reshape expectations across language over the near term.

For readers focused on practical impact, the best next step is to watch what changes around sarvam once attention turns into execution.

Frequently Asked Questions

Why does Sarvam AI Saaras V4 matter right now?

Revolutionizing Speech Recognition for India and BeyondFor readers tracking the shift, Sarvam AI has announced the release of Saaras V4, the latest iteration of its advanced speech recognition model.

What broader change could Sarvam AI Saaras V4 signal?

This groundbreaking development marks a significant leap forward, offering state-of-the-art accuracy across all 22 officially recognized Indian languages, alongside comprehensive support for global English accents.

What should the market watch next around Sarvam AI Saaras V4?

Saaras V4 is poised to transform how businesses and developers interact with spoken language, especially in India’s diverse linguistic landscape.Unmatched Linguistic CoverageMeanwhile, One of Saaras V4’s most compelling features is its unparalleled language support.

Source: https://www.marktechpost.com/2026/09/26/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english/

Share this article

Subscribe

By pressing the Subscribe button, you confirm that you have read our Privacy Policy.

Latest News

More Articles