AI Tools.

Search

automatic speech recognition by nvidia

nemotron-3.5-asr-streaming-0.6b

Nemotron-3.5-ASR-Streaming-0.6B is NVIDIA's 600M-parameter cache-aware streaming speech recognition model trained in NeMo. It targets real-time transcription scenarios where audio arrives incrementally, supporting multiple languages with low-latency chunk-by-chunk decoding.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository nvidia/nemotron-3.5-asr-streaming-0.6b at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
nvidia
Pipeline tag
automatic-speech-recognition
Library
NeMo, Transformers
Framework tags
PyTorch
Weight formats
safetensors, GGUF
License tag
other — read the license file in the repo before relying on it
Language tags
multilingual; English (en), Spanish (es), German (de), French (fr), Italian (it), Arabic (ar), Japanese (ja), Korean (ko), Portuguese (pt), Russian (ru), Hindi (hi), Chinese (zh), Vietnamese (vi), Hebrew (he), Dutch (nl), Czech (cs), Danish (da), Polish (pl), Norwegian (no), Swedish (sv), Thai (th), Turkish (tr), Bulgarian (bg), Greek (el), Estonian (et), Finnish (fi), Croatian (hr), Hungarian (hu), Lithuanian (lt), Latvian (lv), Romanian (ro), Slovak (sk), Ukrainian (uk), Maltese (mt), Slovenian (sl)
Papers cited
arXiv:2312.17279, arXiv:2305.05084
Datasets declared
nvidia/Granary, multilingual_librispeech, fleurs, mozilla-foundation/common_voice_8_0, voxpopuli, europarl
Downloads (HF counter at last fetch)
917,259
Likes (HF counter at last fetch)
1,079
Model card
https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b

Use cases

  • Real-time call-center transcription with streaming audio
  • Live captioning for video conferencing
  • Embedding into latency-sensitive voice assistant pipelines
  • Multilingual ASR where a full Whisper-large is too slow

Pros

  • Designed from the ground up for streaming (cache-aware attention)
  • 0.6B parameters keeps inference latency low on modest hardware
  • Strong community validation with 938 likes and 797K+ downloads
  • NeMo ecosystem support for fine-tuning and deployment

Cons

  • Streaming-optimized architecture trades accuracy for latency vs. offline models
  • NeMo framework adds a non-trivial dependency vs. transformers-based ASR
  • Multilingual coverage not detailed in public documentation
  • Accuracy on accented or noisy speech not benchmarked publicly

Tags

nemosafetensorsggufnemotron3_5_asrfeature-extractiontransformersspeech-recognitioncache-aware ASRautomatic-speech-recognitionstreaming-asrmultilingualspeechaudioFastConformerRNNTParakeetASRpytorchNeMoen