AI Tools.

Search

image text to video models

2 models · ranked by HuggingFace downloads

MiniMax-H3

MiniMax-H3 generates synchronized audio-video clips from text or image prompts, producing output with coherent ambient sound and motion together. It supports multiple input modalities including text-to-video, image-to-video, and video-to-video transformation pipelines. The model ships on Diffusers and uses safetensors checkpoints, making it straightforward to integrate into ComfyUI or custom generation workflows.

3,055,205 ↓ · 4,186 ♡

Minimax-H3-nvfp4-INT4-INT8-Convrot

A mixed-precision quantization of MiniMax-H3 using NVFP4, INT4, and INT8 with convolutional rotation (Convrot), targeting NVIDIA hardware for faster video generation with reduced VRAM. This variant is designed for ComfyUI workflows where full-precision weights are not feasible. The quantization is applied to the base MiniMaxAI/MiniMax-H3 weights.

711,076 ↓ · 197 ♡