Case study
Internal toolVoice Typer
A push-to-talk Windows dictation and translation tool with a fully local speech pipeline option (faster-whisper + Ollama) alongside cloud APIs.
FIG. 001 / ARCHITECTURE
How the system connects
Live model
Problem
Cloud-only dictation tools mean every recording leaves the machine, and a single hotkey for both 'write what I said' and 'turn my rambling into a clean English message' didn't exist in one lightweight Windows tool.
System
A Python tray app captures push-to-talk audio and transcribes it either locally (faster-whisper) or via OpenAI Whisper/Google Gemini, selectable per-provider in config. A translate mode runs the local transcript through an Ollama LLM to strip filler and repetition and re-write it as a clean English message; a live-typing mode streams partial transcripts into the focused field roughly every two seconds as the user speaks; and a separate hotkey translates selected text in place.
FIG. 002 / HOW IT WORKS
How it works
Voice Typer is the more advanced sibling of the public Voice to Text tool: the same "speak and it appears" idea, extended with a fully local speech-to-text and cleanup pipeline for anyone who doesn't want audio leaving their machine, plus live incremental typing so text appears as you speak rather than only after you stop. It ships as a proper Windows installer with a system tray UI for switching providers and hotkeys.
Stack
What it runs on
Related systems
Voice to Text + Instant Translate
Open sourceWindows tray tool: double-tap ALT to dictate anywhere via Whisper, or select text and double-tap CTRL for an instant floating translation.
UBA-BRAIN
Internal toolA shared markdown knowledge base and rule system that lets any AI CLI on any machine run the same client work with the same rules, memory and skills.
AI Product Visuals Pipeline
Internal toolA skill-driven pipeline that turns a client's chat request into an on-brand AI product photo, picking the right account, model and reference automatically.
Want a system like this for your team?
We start by mapping one workflow on a 30-minute call.