From the model card
Fields below are copied from the tags and counters on the HuggingFace repository nvidia/personaplex-7b-v1 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.
- Publisher (HF namespace)
- nvidia
- Pipeline tag
- audio-to-audio
- Weight formats
- safetensors
- License tag
other— read the license file in the repo before relying on it- Lineage
-
- base model kyutai/moshiko-pytorch-bf16
- fine-tune of kyutai/moshiko-pytorch-bf16
- Language tags
- English (en)
- Papers cited
- arXiv:2602.06053, arXiv:2503.04721, arXiv:2110.13900, arXiv:2410.00037
- Downloads (HF counter at last fetch)
- 368,110
- Likes (HF counter at last fetch)
- 2,593
- Model card
- https://huggingface.co/nvidia/personaplex-7b-v1
Use cases
- Real-time voice conversation AI with persona customization
- Voice agent development with simultaneous listen/speak capability
- Research into speech-to-speech LLM interaction
- Prototype for voice-driven customer service agents
Pros
- Full-duplex speech capability — can listen and speak simultaneously
- 7B scale provides reasonable conversational quality
- 2497 likes suggests active community and real usage
- Moshi architecture enables low-latency voice interaction
Cons
- 'Other' license — NVIDIA's commercial terms need verification
- Real-time inference requires substantial GPU compute at 7B
- Moshi-based inference requires specific runtime setup beyond standard Transformers
- No multilingual support — English-primary speech interaction
Tags
moshisafetensorspersonaplexspeech-to-speechagentaudio-to-audioenarxiv:2602.06053arxiv:2503.04721arxiv:2110.13900arxiv:2410.00037base_model:kyutai/moshiko-pytorch-bf16base_model:finetune:kyutai/moshiko-pytorch-bf16license:otherregion:us