- Voice-hermes-pipeline.md: full architecture for SamiType → Hermes direct command - Live transcript overlay with inline correction - Live OS modification via voice (add widgets, change themes, move elements) - Phased rollout plan (POC → correction → context → full integration)
11 KiB
SamiType Voice → Hermes Direct Command Pipeline
The Vision
Speak naturally → see your words transcribed live → correct if needed → Hermes executes.
This is NOT a simple voice command router ("open Firefox"). This is a voice-driven coding/system administration interface where natural language spoken commands go directly to Hermes (the AI agent) for execution — writing code, managing servers, configuring the OS, anything Hermes can do.
The Flow
┌──────────────────────────────────────────────────────────────┐
│ Voice Command Pipeline │
│ │
│ 1. SPEAK 2. SEE 3. CORRECT 4. EXECUTE │
│ │
│ ┌──────┐ ┌────────────┐ ┌──────────┐ ┌──────────┐ │
│ │ Mic │───▶│ Live │───▶│ Inline │───▶│ Hermes │ │
│ │ │ │ Transcript │ │ Edit │ │ API │ │
│ │ │ │ (streaming)│ │ (optional)│ │ │ │
│ └──────┘ └────────────┘ └──────────┘ └────┬─────┘ │
│ │ │
│ ▼ │
│ ┌──────────┐ │
│ │ Result │ │
│ │ (visual │ │
│ │ + TTS) │ │
│ └──────────┘ │
└──────────────────────────────────────────────────────────────┘
Step-by-step
- Speak — Press hotkey or say wake word "Computer". Start talking naturally.
- Live Transcript — Words appear on screen as you speak them, streaming, like live captions. A floating overlay bar at the top or bottom of the screen.
- Correct — When you stop speaking, the transcript stays visible for 2 seconds. If SamiType misheard something:
- Click/tap the word to fix it
- Or press Tab to edit the full text
- Or just say "scratch that" to restart
- Or press Enter immediately to send as-is
- Execute — Corrected text goes to Hermes API as a task. Hermes starts working.
- Result — A notification shows "Hermes is working on: [task]". When done, result appears + optional TTS summary.
Example Sessions
Coding
You: "Computer, create a Python script that monitors disk usage
and sends me a Telegram alert if it goes above 80 percent"
[Live transcript appears word by word]
[2s pause — you see it's correct, press Enter]
Hermes: "Writing disk monitor script..."
Hermes: "Done. Script at /usr/local/bin/disk-monitor.py,
systemd service enabled. Want me to test it?"
System Admin
You: "Computer, check if the triangl... no, the DashCaddy server
is responding"
[Transcript: "check if the triangl... no, the DashCaddy server is responding"]
[You edit "triangl" → delete that part]
Corrected: "check if the DashCaddy server is responding"
[Enter]
Hermes: "DashCaddy is up. Response time 42ms.
Dashboard loads in 1.2s. All containers healthy."
OS Control
You: "Computer, install VLC and set it as default media player"
[Live transcript — correct, Enter]
Hermes: "Installing VLC... done. Set as default for video/x-matroska,
video/mp4, and audio/mpeg."
Technical Architecture
Component 1: SamiType Streaming STT
- Streaming mode: Outputs partial transcriptions every ~200ms
- WebSocket or Unix socket → UI overlay receives chunks
- Confidence scores per word — low-confidence words highlighted in red
- VAD (Voice Activity Detection): Auto-detects when you stop talking → 2s correction window
Component 2: Live Transcript Overlay UI
Lightweight floating widget — options:
Option A: KDE Plasma Widget (native)
- Always-on-top floating panel
- QML-based, renders streaming text
- Click-to-edit inline
- Keyboard shortcuts: Enter=submit, Esc=cancel, Tab=edit mode
- Semi-transparent when idle, solid when active
Option B: Electron/Tauri app (cross-platform)
- Floating window, borderless, always-on-top
- Web tech (React/Svelte) for rapid iteration
- Works on any desktop, not just KDE
Option C: Terminal overlay (tmux/terminal)
- Simplest — a tmux pane or terminal window
- Good for development, not daily use
Recommendation: Start with B (Tauri — lightweight, ~10MB, Rust backend), graduate to A (native KDE widget) when SamiType is mature.
Component 3: Correction Layer
Transcript: "create a python script that monitors [disk] usage"
highlighted = low confidence ^
Click "disk" → text field appears → type "disc" → Enter
Or: ignore, press Enter to send as-is (Hermes will understand context anyway)
- Low-confidence words (< 80% certainty) highlighted in amber/red
- Click any word to edit inline
- Tab → enter full-edit mode (entire text becomes editable)
- Auto-submit timer: 2s after speech ends, auto-sends unless editing
- "Scratch that" → clears and restarts listening
Component 4: Hermes API Bridge
# Voice command → Hermes HTTP API
import requests
def send_to_hermes(transcript: str):
response = requests.post("http://localhost:8765/command", json={
"text": transcript,
"source": "voice",
"context": {
"hostname": "samios",
"user": "sami",
"cwd": os.getcwd(),
"active_app": get_active_window()
}
})
return response.json()
# Hermes processes it as a normal task — same as if Sami typed it
Component 5: Result Display
- Notification: KDE notification "Hermes: [task summary]" with status
- TTS: Optional spoken confirmation using SamiType TTS
- Full output: Click notification → opens Konsole with full output log
Implementation Phases
Phase 1: Proof of Concept (1-2 weeks)
- SamiType running in streaming mode, outputting to stdout/JSON
- Simple Python/Tauri overlay showing live transcript
- Enter key → send to Hermes CLI (
hermes ask "transcript text") - Basic notification on completion
- Milestone: Voice "create a file called test.txt" → Hermes creates it → notification
Phase 2: Correction + Polish (2-3 weeks)
- Low-confidence word highlighting
- Inline word editing
- 2s auto-submit timer with visual countdown
- "Scratch that" voice command to restart
- Result TTS via SamiType
- Milestone: Natural correction flow, reliable daily use
Phase 3: Context Awareness (3-4 weeks)
- Active window/app context passed to Hermes
- "Fix this function" when code editor is focused → Hermes sees the file
- "What's wrong with this page" when browser is focused → Hermes checks URL
- Multi-turn: "Now make it async" (Hermes remembers previous command)
Phase 4: Full Integration (ongoing)
- Native KDE Plasma widget
- Systemd service (always available)
- Wake word "Computer" (openWakeWord)
- Custom command shortcuts ("Computer, deploy" = predefined multi-step task)
- Voice history log + replay
Live OS Modification (Killer Feature)
The most powerful use case: designing and modifying the OS in real-time via voice.
You: "Computer, add a weather widget to the top right corner"
[Enter]
Hermes: "Adding weather widget..."
→ KDE plasma widget installed, positioned, configured
→ Widget appears on your screen 3 seconds later
You: "Move it down a bit"
[Enter]
→ Widget shifts down 50px
You: "Make it show 5 days instead of 3"
[Enter]
→ Widget reconfigures, shows 5-day forecast
How This Works
- Hermes receives voice command with desktop context (active screen, resolution, current layout)
- Hermes generates the config change (Plasma widget config, kwin rule, CSS/SVG theme edit)
- Hermes pushes the change via:
qdbuscalls to Plasma (move/resize/configure widgets)- File writes to
~/.config/(theme changes, panel configs) kstart/kquitappto reload components
- Change appears live — no reboot, no logout, instant visual update
Supported Live Modifications
| Voice Command | Mechanism |
|---|---|
| "Add [widget] to [position]" | Install + plasma-shell script |
| "Move [element] [direction]" | qdbus org.kde.plasmashell |
| "Change wallpaper to [name]" | plasma-apply-wallpaper |
| "Make panel transparent" | Edit panel config → qdbus reload |
| "Change accent color to blue" | plasma-apply-colorscheme |
| "Increase font size" | Edit kdeglobals → reload KWin |
| "Add a system tray for [app]" | Install + configure tray widget |
| "Switch to dark mode" | plasma-apply-colorscheme Win11Dark |
| "Add keyboard shortcut [key] for [action]" | Write kglobalshortcutsrc entry |
The Iteration Loop
Speak → See change → "No, not like that" → Hermes adjusts → "Yes, perfect"
This is the same loop you and I already do over Telegram, but:
- Instant (no typing, no copy-paste)
- Visual (you see the result immediately on your screen)
- Iterative ("move it left", "bigger", "more transparent")
- Contextual (Hermes sees your current desktop state)
Key Design Decisions
Why Hermes API and not local command parsing?
- Hermes understands natural language ("make the thing faster" → identifies what "the thing" is from context)
- Hermes can write code, manage servers, fix bugs — not just launch apps
- Hermes has memory of prior commands and project context
- Simple command parsing can't do "refactor the auth module to use JWT"
Why correction instead of just sending raw transcript?
- SamiType will occasionally mishear technical terms (package names, URLs, code syntax)
- 2-second glance + quick fix prevents cascading errors
- Hermes is smart enough to handle most errors, but "install postgress" vs "install postgresql" matters
- You said you want this — the ability to see and fix in real-time
Why Tauri for the overlay?
- ~10MB binary (vs Electron's ~150MB)
- Rust backend handles audio socket efficiently
- Web frontend (React/Svelte) for rapid UI iteration
- Cross-platform if SamiOS ever runs outside WSL