I had ~60 hours of .ogg and .m4a audio recordings and was keen to see if AI could transcribe them without sending the files to a cloud-based service. I see huge potential in local AI models, both from an environmental and privacy perspective, but this was my first real-world test.

My hardware was a modest ThinkPad T14 Gen 1 AMD (equipped with a Ryzen 5 PRO, 16GB RAM, and no dedicated GPU). Running local AI models without a discrete graphics card can be slow and painful, so the setup needed to be optimised to run entirely on CPU without bringing the machine to its knees.

The underlying pipeline relies on faster-whisper, a re-implementation of OpenAI’s Whisper model that uses CTranslate2 to perform fast inference on CPU using 8-bit quantisation. Wrapped safely inside a lightweight Docker container, the script mounts local directories, handles format conversions via ffmpeg, and outputs clean text files with timestamps.

To keep the pipeline efficient and dependable over time, a few key choices were made:

  • Model selection: large-v3-turbo is the default as the optimal balance between high transcription accuracy and light memory usage on CPU. The model can be easily changed using environment variables.
  • Real-time logging: Python’s console output was set to unbuffered mode (-u) to stream transcribed segments directly to the terminal as they are generated.
  • Silence handling: Voice Activity Detection (vad_filter=True) was enabled to strip dead air before processing, cutting down runtime significantly and preventing hallucinations.

Building the script was a genuine collaborative effort between two AI assistants. Claude 4.6 Sonnet drafted the initial shell wrapper and structured the overall Docker orchestration; Gemini 3.6 Flash then reviewed the execution pipeline, unbuffered stdout flags, .env schema, and VAD filter settings to ensure output streamed properly without file buffer delays.

The project is ready to run. You can find the repository on GitHub at jamesgreenblue/local-audio-transcriber. Changes are welcome so please don’t hesitate to open an issue if you spot a problem, or open a pull request for enhancements.