KaigiAI
A multilingual meeting assistant created to help follow fast conversations and translate them with context.
Role
Creator — Rust / Tauri architecture, audio integration, speech recognition and LLM workflows
View KaigiAI sourceSpeech
Microphone and system-audio capture
Context
LLM-assisted translation
Hybrid
Local and cloud inference options
The problem
Following workplace conversations requires both recognizing speech and keeping up with its meaning. Word-for-word translation alone loses useful context.
The assistant also needed to accommodate different hardware. The available work PC could not run local LLM inference at a satisfactory speed.
Desktop speech pipeline
A Rust core manages audio and inference integrations behind a React / TypeScript interface.
- Local whisper.cpp and llama.cpp sidecars expose OpenAI-compatible localhost APIs.
- Grok and Gemini provide cloud inference options; Grok was used for translation in actual meetings when local compute was insufficient.
- Optional ONNX-based diarization supports speaker processing, with observed limitations.
Engineering decisions
Keep inference replaceable
Support local and cloud providers so hardware constraints do not require rebuilding the application.
Use conversation context
Pass speech-derived text through an LLM layer for meaning-aware translation and analysis.
Test in real meetings
Validate the workflow under actual conversation conditions and identify audio-volume and speaker-recognition issues.
Preserve local processing options
Provide a local-first mode for sensitive audio and transcripts; cloud processing depends on the selected provider.
What was validated
The assistant was used in actual workplace meetings, including a cloud-backed translation path when the work PC was insufficient for local inference.
Observed limitations included inconsistent audio volume and imperfect speaker recognition. There is no formal accuracy benchmark or claim of commercial production-scale deployment.
Open-source project
Public scopeThe source is public. Meeting recordings, transcripts and workplace data are not included in this case study. A server-based inference deployment is a future option rather than a completed production configuration.
Further discussion
- →Improve audio-volume handling
- →Evaluate speaker recognition systematically
- →Assess a server inference option for lower-powered clients
Technology