Hello, I'm Uba

AI Automation Builder

All systems

21 of 24AI systems / 21 of 24

Case study

Internal tool

Voice Typer

A push-to-talk Windows dictation and translation tool with a fully local speech pipeline option (faster-whisper + Ollama) alongside cloud APIs.

Pythonfaster-whisperOllamaOpenAI Whisper APIGoogle Gemini APIWindows global hotkeys

FIG. 001 / ARCHITECTURE

How the system connects

Live model

Problem

Cloud-only dictation tools mean every recording leaves the machine, and a single hotkey for both 'write what I said' and 'turn my rambling into a clean English message' didn't exist in one lightweight Windows tool.

System

A Python tray app captures push-to-talk audio and transcribes it either locally (faster-whisper) or via OpenAI Whisper/Google Gemini, selectable per-provider in config. A translate mode runs the local transcript through an Ollama LLM to strip filler and repetition and re-write it as a clean English message; a live-typing mode streams partial transcripts into the focused field roughly every two seconds as the user speaks; and a separate hotkey translates selected text in place.

FIG. 002 / HOW IT WORKS

How it works

Voice Typer is the more advanced sibling of the public Voice to Text tool: the same "speak and it appears" idea, extended with a fully local speech-to-text and cleanup pipeline for anyone who doesn't want audio leaving their machine, plus live incremental typing so text appears as you speak rather than only after you stop. It ships as a proper Windows installer with a system tray UI for switching providers and hotkeys.

Stack

What it runs on

Pythonfaster-whisperOllamaOpenAI Whisper APIGoogle Gemini APIWindows global hotkeys

Want a system like this for your team?

We start by mapping one workflow on a 30-minute call.

Message me onWhatsAppTelegram