Flinx
Voice dictation for Linux, built to solve a gap in my own daily workflow.
The challenge
I built Flinx while using Fedora Linux 44 with KDE Plasma as my daily driver. Voice dictation had become part of how I worked, but Vowen was not available on that setup. I wanted a simple interaction: keep working in an editor, browser or terminal, hold a shortcut to speak, then release it to insert the text. Building that experience on Wayland meant solving more than speech recognition. A background tool needed a dependable global trigger, a recording indicator that left the active application in control, and a way to deliver text across different compositor capabilities. The product started with a problem I experienced myself and a workflow I wanted to keep.

The system
I designed Flinx as a Python desktop application with a PyQt6 interface and a layered dictation engine. Input detection, microphone capture, transcription and text delivery each have a separate responsibility, coordinated through explicit idle, recording, processing, pasting and error states. The floating pill communicates the current state without becoming a window the user has to manage. A control center and system tray handle configuration and everyday operation. The engineering follows the same interaction throughout: hold, speak, release, continue. Audio capture stays local until the completed recording is submitted to Groq for transcription, while background workers keep that network request and text delivery away from the interface thread.
- 01
Hold to record
A configurable shortcut starts microphone capture from the current application. The default combines left and right Shift. The hotkey engine reads keyboard events through evdev without taking an exclusive device grab, so the keyboard remains available to the application the user is working in.
- 02
See and hear the state
sounddevice captures 16 kHz mono audio into a NumPy buffer. Microphone amplitude drives the pill’s live waveform, alongside a recording indicator and elapsed time. Synthesized start and stop sounds provide another cue without requiring the user to watch the overlay.
- 03
Transcribe and refine
Releasing the shortcut produces a temporary WAV. A background worker sends the completed clip to Groq’s Whisper Large v3 Turbo model using the user’s API key, with configurable language and vocabulary hints. Empty results and common silence artifacts are filtered. An optional Llama pass removes fillers and adjusts punctuation, falling back to the transcript if refinement fails.
- 04
Return text to the application
Flinx checks whether the compositor supports wtype for direct text input. Otherwise, it copies the transcript with wl-copy and uses ydotool to issue the paste shortcut. The pill is hidden before delivery, temporary audio is removed, and the engine returns to its idle state.
Decisions that shaped it
Build around the desktop I used
Fedora KDE was the starting environment, so Linux desktop integration shaped the architecture from the outset. The global shortcut uses evdev rather than relying on an application-specific key listener. Input permissions and the text injection service are handled through setup and system checks, making those operating-system requirements part of the delivered tool.
Make feedback coexist with focus
The pill is a translucent PyQt6 surface configured to avoid activation and pass pointer input through. It shows recording and processing feedback while the user stays in their current task. Hiding it before pasting and allowing focus to settle addresses the handoff between visual feedback and text delivery, where a small interruption would undermine the whole interaction.
Separate interface and background work
Hotkey listening and the transcription-and-paste pipeline run outside the UI thread. Explicit state transitions coordinate the recorder, pill and tray rather than letting each surface invent its own status. Cleanup runs after processing, including error paths, so temporary recordings are removed and the application can recover for the next dictation without a manual restart.
Adapt text delivery to the compositor
Wayland desktops do not all expose the same text input capabilities. Flinx tests for wtype support and falls back to the Wayland clipboard plus ydotool when needed. Keeping that strategy in a dedicated delivery layer lets the dictation workflow remain consistent while the mechanism adapts to the desktop underneath it. This was a central systems decision, not an interface detail.
Keep transcript refinement optional
Dictation should return the user’s words. The optional cleaner is instructed to format text, remove fillers and preserve meaning rather than respond to its contents. It is disabled by default and falls back to the transcript if the extra request fails. Language settings and transcription hints offer further control without making every recording depend on a second model call.
Deliver the surrounding product
The control center includes microphone testing, shortcut configuration, API validation, language settings and system checks. The tray exposes status, the last transcript and configuration reloads. Installation scripts cover Fedora, Arch, Debian/Ubuntu and openSUSE, with user services and desktop launchers. PyInstaller packaging and an optional AppImage build give the project a distribution path beyond running Python source from a terminal.
Delivered
Flinx restored voice dictation to my own Linux workflow and became a complete open-source desktop application: a global trigger, audio capture, live feedback, cloud transcription, optional text refinement and delivery into the active application. I owned the product decisions, interface and operating-system integration together, including the settings, diagnostics, installation and packaging around the core engine. Published under the MIT license, the repository makes those decisions inspectable. The project extends my work beyond the web into Python desktop engineering and Linux systems, while keeping the same standard of care: start with a real problem and carry the solution through to everyday use.