Skip to main content

Migrating from Google Cloud Speech-to-Text to Whisper: Setup Guide

I switched from Google Cloud Speech-to-Text to self-hosted Whisper for local transcription. Here's what changed, what broke, and whether it was worth the work.

I spent two years feeding audio into Google’s Speech-to-Text API. Worked fine. Then the bill started climbing, and I realized I was sending everything — meeting recordings, podcast clips, voice notes — into someone else’s infrastructure. That friction finally pushed me to migrate from Google Cloud Speech-to-Text to Whisper, and I ran the setup yesterday. This is what happened.

Whisper screenshot
Whisper u2014 from the official site

Why I Left Google Cloud Speech-to-Text

The API itself is excellent. Genuinely. It handles accents better than most open-source models, it supports 200+ languages, and the latency is negligible. But I was paying roughly $1.50 per hour of audio transcribed, and my usage kept creeping up. Not because I was running some massive commercial operation—just because it was easy to throw things at it. Meeting recordings. Voicemails I wanted indexed. A few podcasts where I wanted searchable transcripts.

The real problem wasn’t the money. It was the dependency. If Google rate-limits me, my workflow breaks. If they change pricing, I adjust. If they decide the API tier I use isn’t profitable, it gets deprecated. And every transcription generates a log somewhere that I can’t control.

Self-hosted Whisper removes that layer. It’s open-source, it lives on my hardware, and nothing leaves my network. The catch—and there is one—is that accuracy takes a hit on edge cases, and the hardware requirements are real.

Advertisement

Prerequisites and Hardware

I’m running this on a Dell OptiPlex 7090 I keep in the closet. It’s got an i7-10700K, 32GB RAM, and an RTX 3060. For Whisper specifically, I’m using the base model, which sits at about 140MB and runs inference on CPU in around 2–3 seconds per minute of audio. If I had gone with the large model (3.1GB), I’d see better accuracy but slower processing.

You don’t strictly need a GPU. CPU inference works fine if you’re not in a hurry. But if you’re doing transcription on demand—like transcribing meeting recordings within a few seconds—a GPU makes the difference between usable and tolerable.

Memory matters more than I expected. The base` model needs around 4GB of free RAM to avoid thrashing. I have 32GB total, but Docker and my other services (Home Assistant, Jellyfin, a few databases) take up 12–14GB before Whisper even starts. That left me with a comfortable margin. On a tighter system, I'd probably stick to the tiny model and accept lower accuracy.

Getting Whisper Running with Docker and faster-whisper

I didn't use the vanilla OpenAI Whisper package. Instead, I went with faster-whisper, a community fork that optimizes inference significantly. On CPU, it cuts processing time roughly in half. The trade-off is minimal—it's less widely tested, but it's been solid for months in my environment.

My docker-compose setup looks like this:

version: '3.8'

services:
  whisper:
    image: onerahmet/openai-whisper-api:latest
    container_name: whisper
    ports:
      - "8000:8000"
    environment:
      - WHISPER_MODEL=base
      - DEVICE=cuda
    volumes:
      - ./audio:/audio
      - ./transcripts:/transcripts
    restart: unless-stopped
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

The first run pulled the model and took about 5 minutes. Now it's warm, and API calls respond in well under a second for small files.

The Accuracy Gap (and Why It Matters)

Here's where I hit real friction. Google Cloud Speech-to-Text got 97% of what I threw at it correct. Whisper with the base model gets roughly 93% in English, and that percentage drops faster with background noise, accents outside North American English, or technical jargon.

For my use case, that's usually fine. Meeting notes need 90% accuracy—you can skim and catch the misses. Podcast transcripts are the same way. But I have a handful of calls where getting every single word right matters, and for those, I'm still using Google's API as a fallback. It's not ideal, but it's honest.

The larger models (small, medium, large) close that gap, but they're correspondingly slower and heavier. A two-hour meeting takes 20+ minutes to process with the large model on my CPU. That's too slow for interactive use.

Integration with Home Assistant and Local Workflows

Once Whisper was running, I wired it into Home Assistant for voice command transcription. Instead of sending voice clips to Google, they now hit my local Whisper API on port 8000. The latency is maybe 100ms longer than cloud, but it's all local, and that feels different—cleaner, somehow.

I also wrote a simple Python script that watches a folder, transcribes anything new, and dumps the output as text files. Useful for batch jobs. Nothing fancy, but it saves me from manually hitting the API every time.

#!/usr/bin/env python3
import requests
import os
from pathlib import Path

audio_dir = Path('/mnt/audio_in')
output_dir = Path('/mnt/transcripts')

for audio_file in audio_dir.glob('*.mp3'):
    with open(audio_file, 'rb') as f:
        files = {'file': f}
        resp = requests.post('http://localhost:8000/asr', files=files)
        transcript = resp.json()['text']
        
    output_file = output_dir / f"{audio_file.stem}.txt"
    output_file.write_text(transcript)
    print(f"Transcribed {audio_file.name}")

The Home Assistant integration also means I can trigger transcription from automations. If someone leaves a voice note on a specific channel, it automatically gets transcribed and posted to a webhook. Useful for accessibility stuff, or just for my own searchability.

What I Miss from Google Cloud

Honestly? Speed. Google's API transcribes in near real-time. Whisper, even on GPU, needs a few seconds. For a 30-minute meeting, that's acceptable. For voice commands in Home Assistant, you start to feel the lag if you're used to cloud-fast.

Support is another one. When something breaks with Google, I open a ticket. When something breaks with Whisper, I read GitHub issues and try to debug it myself. That's the self-hosted trade-off, and I knew it going in, but it's worth naming.

And there's the accuracy thing I mentioned. Google handles Spanglish, heavy accents, and technical terminology better. I've worked around it by chunking audio differently—breaking longer files into smaller segments—but it adds complexity.

The Cost Reality

I'm no longer paying Google, which is the obvious win. But my homelab is running 24/7 anyway, and Whisper adds maybe 15–20W under load, which is rounding error on my electric bill. The upfront effort—learning Docker, setting up the API wrapper, writing integration code—that was the real cost.

If you're already running a homelab, Whisper's financial cost is basically zero. If you're not, you're now running a whole server just for transcription, and that math changes significantly. The decision isn't really about money. It's about control.

Common Pitfalls

I made a few mistakes the first run. First, I forgot to allocate GPU memory correctly in Docker, so it kept falling back to CPU inference. The container logs gave no indication of this—I only noticed because throughput was abysmal. Second, I didn't account for the model download size. The base model is 140MB, the small is 500MB, and medium is over a gig. Knowing those numbers upfront would have saved me wondering why the container startup took forever.

Third, and this is embarrassing, I assumed Whisper could ingest any audio format. It can't. MP3s need to be decoded first. I wrote a quick FFmpeg wrapper to handle conversion automatically, but it's an extra step if you're not careful.

FAQ

How accurate is Whisper compared to Google Cloud Speech-to-Text?

Whisper's base model runs about 93–95% accurate in English on clean audio. Google Cloud hits 97%+. The gap widens with accents, background noise, and technical terms. For most hobby use, Whisper is fine; for professional transcription, Google is still more reliable.

Can Whisper run on a Raspberry Pi?

Technically yes, but you'll want a Pi 4 with at least 4GB RAM and the tiny model. Processing will be slow—5+ minutes per hour of audio. It works, but it's not practical for interactive use.

What's the difference between Whisper and faster-whisper?

Faster-whisper is a community optimization of OpenAI's model. It speeds up inference by roughly 50% on CPU with minimal accuracy loss. If you care about speed on limited hardware, it's worth using.

Do I need a GPU to run Whisper?

No. CPU inference works fine, just slower. For a two-hour meeting, CPU might take 5–10 minutes; GPU takes 1–2. Depends on your patience and hardware.

How much bandwidth does Whisper use?

Zero outside your network. Audio stays local. This is the whole point—nothing leaves your homelab.

Explore Whisper in our AI Homelab Toolkit.

Written by Engineer running a 24/7 homelab since 2022. Every guide here is built and tested on my own hardware. No paid placements.

Keep reading