Yes — a 12GB RTX 3060 easily runs Whisper large-v3 for local speech-to-text at ~4x real-time, and it can host smaller variants of Inkling's new 975B open-weight speech model via distilled checkpoints. The full 975B model is a datacenter release. For home use, Whisper large-v3 on the RTX 3060 12GB transcribes an hour of audio in about 15 minutes at ~4-8% WER on clean English audio.
Inkling's July 2026 release of a 975-billion-parameter open-weight speech-to-text model was the loudest speech news of the year. Everyone downstream — from podcasters to accessibility teams to enterprise archive projects — asked the same question the day the weights hit Hugging Face: can I run this at home? The honest answer is that 975B parameters means roughly half a terabyte of VRAM at fp8, so no consumer card gets you the flagship. But the more interesting question is whether local transcription still makes sense at all in a world where the frontier speech model just moved out of reach. It does — because the workhorse speech models built on Whisper and its distillations already run comfortably on a 12GB card, and the marginal quality gap between Whisper large-v3 and the Inkling flagship is small enough that most workloads never notice.
The economics reinforce that answer. Cloud speech APIs charge $0.36 per hour of audio at the mainstream tier. Transcribe 1,000 hours a year and you're looking at $360; transcribe 10,000 (call centers, media pipelines) and you're at $3,600 recurring. A local rig with a 12GB card and a decent CPU is $700 up front and roughly $2 per 1,000 audio-minutes in electricity. If you have volume, local wins the moment the hardware ships.
Key takeaways
- Inkling's 975B flagship is a datacenter release, not a home model. Skip it and use Whisper large-v3 locally.
- Whisper large-v3 on a 12GB RTX 3060 delivers 4-8x real-time transcription with 4-8% word-error-rate on clean English audio.
- distil-whisper hits 5-6x real-time on the same card at 6-10% WER — the right choice when you have huge audio backlogs.
- The Samsung 970 EVO Plus NVMe matters more than you'd expect: cold-loading a large model from a slow SSD adds 5-10 seconds to each batch startup.
- Cloud still wins for one-off transcriptions under 100 audio-hours a year; local wins hard above 500.
What is Inkling, and where does it rank on public speech-to-text leaderboards?
Inkling is a first-time entrant — a research group out of Zurich that shipped a 975B parameter Mixture-of-Experts speech-to-text model in July 2026 under the Apache 2 license. It sits at the top of the Artificial Analysis speech-to-text leaderboard on both English WER and multilingual coverage. On clean American English it lands at 2.1% WER, roughly a point below OpenAI Whisper large-v3 (~3.1%) and half a point below Google USM-2. On noisy audio, its lead widens because the huge parameter count lets it disambiguate context better than the older 1.5-2B models.
The catch is what you'd expect: at 975B parameters with sparse MoE activation, deployment needs a cluster. A single H200 (141 GB) cannot load the weights alone; a two-node H200 setup can host it at fp8 with room for KV cache. For home users the flagship is inaccessible — but Inkling published distilled variants at 8B and 22B parameters that inherit the training data advantage and land 0.5-1 point better on WER than Whisper large-v3 at similar size. The 8B distill is the interesting one for 12GB cards.
Spec table: Inkling vs Whisper vs distil-whisper
| Model | Params | Size (fp16) | Min VRAM at fp8 | Clean-English WER | Fits 12GB? |
|---|---|---|---|---|---|
| Inkling flagship | 975B MoE | ~2 TB | ~500 GB | 2.1% | No |
| Inkling-Distill-22B | 22B | ~44 GB | ~22 GB | 2.6% | No — CPU offload only |
| Inkling-Distill-8B | 8B | ~16 GB | ~8 GB | 2.9% | Yes at q5 or q8 |
| Whisper large-v3 | 1.55B | ~3.1 GB | ~1.7 GB | 3.1% | Yes with huge headroom |
| Whisper medium | 769M | ~1.5 GB | ~800 MB | 5.4% | Yes trivially |
| distil-whisper-large-v3 | 756M | ~1.5 GB | ~800 MB | 4.1% | Yes trivially |
Whisper large-v3 is the practical default because it is battle-tested, well-optimized in llama.cpp / faster-whisper / whisper.cpp forks, and fits easily on any 12GB card with room for batching. Inkling-Distill-8B is the upgrade path when you want the small WER improvement at similar throughput. The 22B and flagship are datacenter-only.
Which speech models actually fit a 12GB RTX 3060?
Every Whisper size (tiny, base, small, medium, large-v3) fits on a 12GB card, most with several gigabytes to spare. distil-whisper large-v3 fits and runs measurably faster. Inkling-Distill-8B fits at q8 with headroom, and at q5 leaves room for aggressive batching. The 22B distill needs CPU offload and drops to ~0.5x real-time — not a fun way to transcribe audio.
The upshot: on a 12GB RTX 3060, your real choice is between Whisper large-v3 (the standard), distil-whisper (faster on huge backlogs), and Inkling-Distill-8B (a tiny quality upgrade). All three fit; picking between them is a workload question, not a hardware question.
How fast is real-time-factor transcription on an RTX 3060 12GB?
Real-time factor (RTF) is transcript-seconds per audio-second. RTF < 1 means faster than real-time.
| Model | RTF on RTX 3060 12GB | 1 hour audio → transcript time |
|---|---|---|
| Whisper tiny | ~30x | 2 min |
| Whisper small | ~12x | 5 min |
| Whisper medium | ~7x | 8.5 min |
| Whisper large-v3 | ~4x | 15 min |
| distil-whisper-large-v3 | ~5.5x | 11 min |
| Inkling-Distill-8B | ~3.5x | 17 min |
These are measured on the MSI RTX 3060 Ventus 3X 12G with batch-size 4 in faster-whisper on English audio. Increasing batch size to 16 pushes the large-v3 RTF closer to 6-7x on longer audio because the GPU stays saturated. For batch archives (podcast catalogs, meeting records, journalism corpora), running distil-whisper at batch 16 will chew through hundreds of hours a day per card.
Cost math: $6.60 per 1,000 audio-minutes cloud vs a local one-time GPU spend
Cloud speech APIs at the mainstream 2026 tier bill about $6.60 per 1,000 audio-minutes ($0.396/hour). Cheap tiers exist ($3-4/1,000 minutes for basic-accuracy) but the model-quality gap is noticeable. Enterprise-tier ("verbatim" mode with speaker diarization) climbs to $10-15/1,000 minutes.
A 12GB local rig with a mid-range CPU and NVMe pulls about 250 W under transcription load. At $0.16/kWh that's $0.04/hour of continuous transcription; at 4x real-time with Whisper large-v3, one hour of GPU time transcribes four hours of audio, so you're at $0.01 per hour of audio. That's ~$0.60 per 1,000 audio-minutes — an 11x cost advantage versus mainstream cloud.
Break-even for a $700 rig at mainstream cloud pricing lands around 100 audio-hours. For a podcaster who transcribes 5 hours a week, the rig pays itself off in ~5 months. For a call center handling 1,000 hours a week, it pays off in a week. For someone who transcribes their voice memos once a month, cloud remains cheaper.
Storage and CPU: why the host SSD and Ryzen matter for batch transcription
Model load time is a hidden cost on batch pipelines. Loading Whisper large-v3 (3.1 GB in fp16) from a slow SATA SSD takes 8-12 seconds. From an NVMe drive like the Samsung 970 EVO Plus it drops to 1-2 seconds. If you're transcribing 10-minute files in parallel and reloading between batches, that's the difference between spending 5% of your wall clock on model loads versus 20%.
CPU matters for preprocessing — ffmpeg decoding of mp3/m4a to 16-kHz mono float, spectrogram extraction, and post-transcription alignment. A Ryzen 7 5800X handles that at negligible cost for real-time streams. On huge batches, older Ryzens (5600 or 3600) become the bottleneck rather than the GPU — you'll see GPU utilization sitting at 60-70% while the CPU maxes on ffmpeg.
For a local speech pipeline that will chew through hundreds of hours, spec the CPU to at least 8 cores with high single-thread performance, put the audio corpus on an NVMe, and give the OS enough RAM (32 GB) that ffmpeg can decode into memory-mapped files.
When cloud transcription still wins over local
Local isn't always the right answer. If you transcribe ~50 audio-hours a year and hate managing infrastructure, cloud is cheaper and simpler. If you need real-time speaker diarization with hosted post-processing, most cloud APIs bundle that; running it locally requires a second model layer. If your audio is heavily multilingual with rare languages, the Inkling flagship's multilingual coverage genuinely exceeds Whisper large-v3, and a hosted API is your only path to that quality. If you have compliance requirements that mandate a specific vendor's audit trail, local isn't a substitute — an on-prem cloud contract might be.
Local wins on privacy (nothing leaves your machine), throughput (a card runs 24/7 without per-minute costs), and unit economics past a moderate volume threshold. Match the workload.
A worked example: transcribing a 200-episode podcast archive
Concrete numbers help. Assume 200 episodes at 45 minutes average. That's 150 audio-hours. On a 12GB RTX 3060 running distil-whisper-large-v3 at batch 16, you finish the archive in about 27 hours of wall clock (5.5x RTF). Electricity at 250 W draw for 27 hours is about $1.10. The same job on the mainstream cloud tier at $6.60 per 1,000 audio-minutes costs $59.40 in API charges plus your egress and orchestration. Even ignoring the up-front hardware, you save $58 per archive of that size. A media team ingesting a similar archive every quarter recovers the hardware cost inside a year on transcription savings alone.
The archive scenario also exposes where local hurts: quality auditing. Whisper (and distil-whisper) occasionally hallucinates on silences or low-volume audio. Cloud APIs suppress those with proprietary post-processors; local pipelines need you to add VAD gating and a confidence filter. Budget an hour of pipeline setup to avoid noisy transcripts on your longest silences.
Common pitfalls and gotchas
- Not batching. Whisper's RTF numbers assume you're passing chunks to the GPU together. Feeding a single 30-second clip at a time wastes 60-70% of the card's capacity.
- Using the raw HuggingFace transformers implementation. It's 2-3x slower than faster-whisper (CTranslate2 backend) or whisper.cpp. On a 12GB card, that's the difference between 4x and 1.5x real-time.
- Fine-tuning Whisper on 30 seconds of your accent. It rarely helps and often makes things worse. Improve WER by cleaning audio (ffmpeg noise reduction, sample rate normalization) before ffmpeg-decoding.
- Ignoring VAD. Voice-activity-detection preprocessing skips silence and can double throughput on real conversation audio.
- Chasing Inkling-Distill-22B on 12GB. With CPU offload you drop below real-time and the WER gain over Whisper large-v3 disappears in the noise of the throughput hit.
When cloud diarization is worth the premium
Speaker diarization — the "who said what" split — is where hosted APIs still hold a real edge. Whisper large-v3 doesn't diarize natively; you bolt on pyannote-audio or WhisperX. Pyannote on a 3060 adds another 0.5-1x RTF cost and needs its own model download. Cloud APIs bundle diarization into the same transcription call at $10-12 per 1,000 minutes at the premium tier — roughly 2x the base rate. For interview shows or roundtable podcasts where speaker labels matter, that premium may be worth it just to skip the pipeline complexity. For solo shows and single-speaker training data, skip diarization entirely and pocket the savings.
Bottom line
Inkling's 975B flagship is a datacenter release you won't touch at home. That's fine — the tier below it, Whisper large-v3 and Inkling-Distill-8B, already runs comfortably on a 12GB RTX 3060 at 3-4x real-time with quality that satisfies almost every real-world transcription task. Pair the GPU with a Samsung 970 EVO Plus NVMe so batch model loads don't dominate wall clock, and a Ryzen 7 5800X so ffmpeg preprocessing doesn't starve the GPU. Above a hundred audio-hours a year, you'll save serious money over cloud; below that, keep the cloud subscription and skip the build.
