DL
Selected work
Open-source speech AI case study

KaigiAI

A multilingual meeting assistant created to help follow fast conversations and translate them with context.

Personal open-source project — used in actual work meetings

Role

Creator — Rust / Tauri architecture, audio integration, speech recognition and LLM workflows

View KaigiAI source

Speech

Microphone and system-audio capture

Context

LLM-assisted translation

Hybrid

Local and cloud inference options

The problem

Following workplace conversations requires both recognizing speech and keeping up with its meaning. Word-for-word translation alone loses useful context.

The assistant also needed to accommodate different hardware. The available work PC could not run local LLM inference at a satisfactory speed.

Desktop speech pipeline

A Rust core manages audio and inference integrations behind a React / TypeScript interface.

01WASAPI audio
02whisper.cpp recognition
03Transcript
04Local or cloud LLM
05Contextual translation
06SQLite history
  • Local whisper.cpp and llama.cpp sidecars expose OpenAI-compatible localhost APIs.
  • Grok and Gemini provide cloud inference options; Grok was used for translation in actual meetings when local compute was insufficient.
  • Optional ONNX-based diarization supports speaker processing, with observed limitations.

Engineering decisions

01

Keep inference replaceable

Support local and cloud providers so hardware constraints do not require rebuilding the application.

02

Use conversation context

Pass speech-derived text through an LLM layer for meaning-aware translation and analysis.

03

Test in real meetings

Validate the workflow under actual conversation conditions and identify audio-volume and speaker-recognition issues.

04

Preserve local processing options

Provide a local-first mode for sensitive audio and transcripts; cloud processing depends on the selected provider.

What was validated

The assistant was used in actual workplace meetings, including a cloud-backed translation path when the work PC was insufficient for local inference.

Observed limitations included inconsistent audio volume and imperfect speaker recognition. There is no formal accuracy benchmark or claim of commercial production-scale deployment.

Open-source project

Public scope

The source is public. Meeting recordings, transcripts and workplace data are not included in this case study. A server-based inference deployment is a future option rather than a completed production configuration.

Further discussion

  • →Improve audio-volume handling
  • →Evaluate speaker recognition systematically
  • →Assess a server inference option for lower-powered clients

Technology

RustTauri v2ReactTypeScriptwhisper.cppllama.cppSQLiteWASAPIGrokGemini