# Voxtral Realtime — weights, license and what is actually open

A compact streaming speech-recognition model designed for low-latency transcription through a realtime API.

Canonical: https://brightaifuture.com/open-models/voxtral
Format: model-family
Source publication: Not established
Bright publication: 2026-09-19
Substantive update: None recorded
Evidence and review: AI-assisted primary-source check on 2026-09-19; no independent model reproduction or license certification. vLLM integration corroborates production serving support; transcription quality claims are not independently reproduced here.

## Documented release

Voxtral-Mini-4B-Realtime-2602

Organization: Mistral AI

Voxtral Realtime release: 2026-02-04 (day precision). Source: https://mistral.ai/news/voxtral-transcribe-2/

Parameters: 4B

Architecture: Streaming audio encoder-decoder model integrated with Mistral inference runtimes.

Modalities: audio, text

Context: Streaming sessions; a comparable language-token window is not asserted here.

## License and commercial use

Apache-2.0

Permitted by Apache-2.0.

## What is open

Weights: available. Official weights are downloadable. Source: https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602

Architecture: available. Configuration and model card are public. Source: https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602

Inference code: available. The release documents vLLM’s Realtime API integration. Source: https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602

Training code: not-established. Complete training code was not confirmed. Source: https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602

Training recipe: not-established. A full reproducible training recipe was not confirmed. Source: https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602

Data information: partial. Capabilities and languages are described without a complete corpus inventory. Source: https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602

Training data: not-established. No sufficiently specific public artifact was confirmed in this review. Source: https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602

Evaluation: partial. Vendor speech evaluations are published in the model materials. Source: https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602

Commercial use: available. The official checkpoint is Apache-2.0 licensed. Source: https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602

## Uses and strengths described in sources

Streaming transcription

Low-latency audio processing

Compact size

## Hardware and quantization

The 4B size supports comparatively modest deployment, but streaming overhead and audio session length affect requirements.

Runtime-supported formats should be verified against the official card

## Independent evidence

vLLM integration corroborates production serving support; transcription quality claims are not independently reproduced here.

## Limitations

Speech-focused rather than a general multimodal assistant

Training data is not fully disclosed

## Provenance and history

{}

## Original sources

- [Voxtral Transcribe 2 and Realtime release](https://mistral.ai/news/voxtral-transcribe-2/)
- [Voxtral Mini 4B Realtime model card](https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602)
- [Mistral model changelog](https://docs.mistral.ai/resources/changelogs)

## Continue exploring

- [Open Models](https://brightaifuture.com/open-models)
- [Open intelligence thread](https://brightaifuture.com/threads/open)
