AI Tools.

Search

audio text to text models

2 models · ranked by HuggingFace downloads

VibeVoice-ASR-HF

by microsoft

VibeVoice-ASR is Microsoft's HuggingFace-packaged automatic speech recognition model, likely a Whisper-style or custom encoder-decoder ASR system targeting informal or conversational speech. The 'Vibe' branding suggests orientation toward natural conversational audio.

528,090 ↓ · 157 ♡

Qwen2-Audio-7B-Instruct

by Qwen

Qwen2-Audio-7B-Instruct is Alibaba's multimodal model handling audio and text inputs, capable of audio analysis, speech-to-text transcription, and audio-grounded Q&A. It's instruction-tuned for dialog about audio content. Apache-2.0 licensed and compatible with the Transformers qwen2_audio model type.

377,338 ↓ · 551 ♡