A local, real-time voice agent you can talk over.
Parakeet listens · Claude Haiku thinks · Kokoro speaks · SQLite remembers
A fully local, real-time voice assistant for Windows — sub-second speech-to-speech latency, barge-in interruption, web search, tool use via MCP, and voice cloning. Built from the ground up with Parakeet (ASR), Silero VAD, Kokoro-82M (TTS), and LangGraph.
| Stage | Latency |
|---|---|
| Parakeet one-pass, 4.4 s clip (hybrid) | 45 ms best / 79 ms median |
| Parakeet one-pass, 64 s monologue (hybrid) | 147 ms best / 169 ms median (RTFx 437) |
| Silero VAD per 32 ms chunk (CPU) | 0.6 ms p50 / 0.7 ms p95 |
| Kokoro first audio from token stream | ~450 ms |
| Groq gpt-oss-20b: first speakable sentence | 270–330 ms |
| Anthropic claude-haiku-4-5: turn end → first audio | 1.7–1.9 s |
| Groq gpt-oss-20b: turn end → first audio | ~1.1 s |
Full benchmark write-up: results/REPORT.md
mic ─ Silero VAD ─┬─ EchoGate ─ StateMachine ─ barge-in
├─ live partials (Parakeet re-decodes every 0.6 s for the UI)
└─ utterance ─ Parakeet ─ Brain (LangGraph + tools, streaming) ─ Kokoro ─ speaker
Full component map: ARCHITECTURE.md
- Real-time transcription — Parakeet TDT 0.6B on CUDA + int8 CPU decoder hybrid; 45–170 ms one-pass latency
- Smart endpointing — Smart Turn v3.2 holds mid-sentence pauses (0.3–0.6 s) and fires at real turn ends
- Two-stage barge-in — 96 ms of speech ducks TTS to 25%; 320 ms of sustained speech cuts it entirely; echo-gated so the agent doesn't interrupt itself
- Streaming TTS pipeline — Kokoro splits LLM output into sentences; first audio in ~450 ms from turn end
- Voice cloning — NeuTTS-Air engine clones any voice from a 3–15 s clip, hot-swappable at runtime
- Web search — SearXNG (self-hosted, no Docker needed) + Trafilatura for full-page reading
- MCP tool support — attach any stdio or HTTP MCP server at runtime; ships with a local demo server
- SQLite memory — full-text search over conversation history; the agent remembers across sessions
- Textual TUI — scrollable conversation, multi-line editor, paste presets, mic mute, F2 wake word
- Groq key pool — rotate up to N keys on rate limits so the conversation never stalls
- Windows 10/11 (tested on Windows 11)
- Python 3.11
- NVIDIA GPU with CUDA 12.x support (driver ≥ 551; driver ≥ 580 lifts all version pins)
- uv package manager
- At least one LLM provider API key (Anthropic, Groq, or Gemini)
CPU-only? The agent falls back to CPU automatically; ASR latency rises to ~300 ms per utterance on a modern CPU.
git clone https://github.com/nodeblackbox/voiceAgent.git
cd voiceAgent
uv venv --python 3.11 .venvuv pip install --python .venv/Scripts/python.exe -r requirements.txt# ORT 1.26 targets CUDA 12; upgrade your driver to ≥ 580 to use newer ORT builds
uv pip install --python .venv/Scripts/python.exe `
"onnxruntime-gpu==1.26.0" `
"nvidia-cudnn-cu12==9.1.1.17" `
"nvidia-cuda-runtime-cu12==12.4.*" `
"nvidia-cublas-cu12==12.4.*" `
"nvidia-cufft-cu12==11.2.*" `
"nvidia-curand-cu12==10.3.9.*" `
"nvidia-cuda-nvrtc-cu12==12.4.*"Why pin cuDNN 9.1? cuDNN ≥ 9.25 fails on drivers below 580 with
CUDNN_BACKEND_API_FAILED, and onnxruntime silently falls back to CPU. The pin keeps it loud and correct.
# PyPI ships a CPU-only torch on Windows; install the CUDA wheel first
pip install torch --index-url https://download.pytorch.org/whl/cu124
uv pip install --python .venv/Scripts/python.exe kokoro==0.9.4copy .env.example .env
# then edit .env and fill in your keysSee .env.example for all supported variables. At minimum you need one of: ANTHROPIC_API_KEY, GROQ_API_KEY, or GEMINI_API_KEY.
Models auto-download from Hugging Face on first run (~4–5 GB total into ~/.cache/huggingface).
# Default: Claude Haiku 4.5, Kokoro af_heart voice, first available mic
.venv\Scripts\python.exe agent\talk.py
# Faster first sentence (Groq inference)
.venv\Scripts\python.exe agent\talk.py --model groq:openai/gpt-oss-20b
# Typed input only (no mic), agent still speaks
.venv\Scripts\python.exe agent\talk.py --no-mic
# Plain log output instead of Textual TUI
.venv\Scripts\python.exe agent\talk.py --plain
# Send turns from the command line (automated testing)
.venv\Scripts\python.exe agent\talk.py --plain --say "what time is it" "roll two dice"| Command | Effect |
|---|---|
/model groq:openai/gpt-oss-20b |
Hot-swap the LLM (fast, smart, opus, sonnet, gemini, qwen are aliases) |
/voice af_bella |
Switch Kokoro voice without reloading |
/tts neutts / /tts kokoro |
Hot-swap TTS engine |
/clone my.wav|exact transcript |
Clone a voice from an audio clip |
/search on|off |
Enable/disable web tools; on starts SearXNG automatically |
/mcp on demo / /mcp off demo |
Attach or detach an MCP server at runtime |
/tools |
List all tools and which server they come from |
/memory notes |
Show memory notes |
/mic |
Toggle mic mute (also: F2 key or the button in the TUI) |
/verbose on|off |
Show/hide internal telemetry in the conversation view |
/help |
Full command list |
Kokoro ships fixed voices (af_heart, af_bella, af_sky, am_adam, etc.). NeuTTS-Air adds zero-shot voice cloning from any 3–15 s clip. It runs in a separate venv because it requires torch ≥ 2.8.
# Set up the NeuTTS venv (one-time, ~5 GB)
uv venv --python 3.11 .venv-neutts
# follow tts/README_NEUTTS.md for the full pin list
# Start the agent on the cloned voice
.venv\Scripts\python.exe agent\talk.py --tts-engine neutts --neutts-voice paulDetails and timing numbers: tts/README_NEUTTS.md
SearXNG runs from source (no Docker required) in its own venv. search/settings.yml enables the JSON API, disables the bot-limiter for localhost, and selects engines that work without a browser (Google, Brave, Startpage, Wikipedia).
.venv\Scripts\python.exe search\run_searxng.py # start
.venv\Scripts\python.exe search\run_searxng.py --check # health checkThe agent auto-starts SearXNG when /search on is used. Tools: web_search, read_page, research.
agent/mcp.json declares servers. A local demo server ships enabled by default (calculator, unit conversion, dice). Add any stdio or streamable-HTTP MCP server:
{
"servers": {
"my-server": {
"transport": "stdio",
"command": "npx",
"args": ["-y", "my-mcp-package"],
"enabled": true,
"description": "..."
}
}
}Attach/detach at runtime with /mcp on my-server / /mcp off my-server.
agent/ Core agent loop, brain, state machine, UI, tools
talk.py Entry point — the full voice loop
brain.py LangGraph agent with tool support
tui.py Textual TUI
endpoint.py Smart Turn endpointing
echo_gate.py Echo cancellation (self-interruption guard)
memory.py SQLite memory with FTS5 search
tools_web.py Web search and page-reading tools
mcp.json MCP server declarations
bench/ Benchmarking scripts (offline ASR, VAD, turn model)
llm/ LLM provider abstraction (LiteLLM, Groq key pool)
tts/ Kokoro and NeuTTS-Air streaming wrappers
search/ SearXNG launcher and settings
models/ Smart Turn ONNX model (v3.2)
audio/ Reference transcripts for offline benchmarks
results/ Benchmark reports and logs
- No streaming ASR — Parakeet runs one-pass after each utterance; streaming (Sherpa-ONNX) is possible but not implemented
- No word boosting — onnx-asr has no context-biasing hook; NeMo on Linux is needed for that
- Windows-only for now — SearXNG has one Unix-only import that is patched; NeuTTS subprocess comms use binary stdin/stdout. PRs for Linux/macOS welcome
MIT
