OpenDictate Logo
OpenDictatev0.2
Models Catalog

On-Device Speech Models

All 100% offline neural models supported by OpenDictate. Zero cloud latency.

Parakeet TDT 110M (int8)

Defaultnon-streaming

Default out-of-the-box model. Ultra-fast with a tiny memory footprint.

Size: 104.3 MBLatency: 78msParams: 110MHardware: CPU / Apple Silicon / GPU

FastConformer 80ms (Streaming)

streaming

Ultra-low 80ms chunk streaming. Words appear live as you speak.

Size: 125 MBLatency: 80msParams: 115MHardware: CPU / GPU

Parakeet Unified 0.6B (Streaming)

streaming

560ms streaming chunks with high acoustic robustness for noisy environments.

Size: 501.4 MBLatency: 560msParams: 600MHardware: Apple Silicon (M1+) / NVIDIA GPU

Parakeet Unified En 0.6B

non-streaming

High-accuracy ASR optimized for complex paragraphs and developer vocabularies.

Size: 501.4 MBLatency: 120msParams: 600MHardware: Apple Silicon / NVIDIA GPU / CPU

Parakeet TDT 0.6B (Multilingual)

non-streaming

Multilingual transcription with fast token duration modeling.

Size: 487.2 MBLatency: 140msParams: 600MHardware: Apple Silicon / NVIDIA GPU

Whisper Turbo (Large v3)

non-streaming

Highest accuracy with flawless casing, punctuation, and markdown formatting.

Size: 563.8 MBLatency: 220msParams: 809MHardware: Apple Silicon / NVIDIA GPU

Whisper Tiny (en)

non-streaming

Lightweight Whisper variant capable of fast execution on minimal CPU hardware.

Size: 118.1 MBLatency: 95msParams: 39MHardware: Any CPU / Laptop

Whisper Base (en)

non-streaming

Balanced speed and accuracy for everyday dictation.

Size: 208.6 MBLatency: 140msParams: 74MHardware: Standard CPU / Laptop

Whisper Small (en)

non-streaming

High-precision transcription for specialized terminology.

Size: 635.7 MBLatency: 280msParams: 244MHardware: GPU or Apple Silicon

Whisper Medium (en)

non-streaming

Deep neural accuracy for long technical transcripts and meetings.

Size: 1905.9 MBLatency: 450msParams: 769MHardware: GPU (4GB+ VRAM) / 16GB+ RAM

Zipformer EN 20M (Live Captions)

utility

Internal live-caption engine generating partial words live while accuracy models compute.

Size: 29 MBLatency: 45msParams: 20MHardware: Any CPU (0.2% load)

Silero VAD v4

utility

Ultra-fast neural speech boundary detector that automatically triggers dictation.

Size: 1.7 MBLatency: 10msParams: 1.2MHardware: Zero-overhead background daemon

Zipformer 3.3M (Wake Word)

utility

Handsfree wake word spotting engine for voice-activated triggering.

Size: 15.8 MBLatency: 20msParams: 3.3MHardware: Any CPU

OpenDictate automatically downloads and manages models on demand when selected in the app.