speech-core — C++17 on-device voice-agent runtime (VAD + STT + diarization + TTS), dual ONNX / LiteRT backends, Apache 2.0
Open-sourcing a C++17 runtime I've been using for real-time voice agents. The orchestration core is pure C++17 with zero ML dependencies (state machine, VAD-driven turn detection, interruption handling, speech queue with cancel/resume, audio utilities — ring buffer, resampler, mel/STFT). Models sit behind small abstract interfaces (STTInterface, TTSInterface, VADInterface, etc.) so consumers can bring their own backend.
Two reference backends, each independently buildable via a CMake option:
- SPEECHCOREWITHONNX → ONNX Runtime (Silero VAD, Parakeet STT, Kokoro TTS, DeepFilterNet3)
- SPEECHCOREWITHLITERT → LiteRT — libLiteRt from Google's ai-edge-litert PyPI wheel (Silero VAD, Parakeet STT, Nemotron streaming STT, Omnilingual STT, Pyannote diarization, WeSpeaker embeddings, VoxCPM2 TTS)
Both CPU today; an optional CUDA / TensorRT execution provider just landed on the ONNX path (SPEECHCOREWITHCUDA, gated, default off, with build-flag + env (SPEECHCOREORTPROVIDER) + runtime-probe resolution and silent CPU fallback).
Platforms: Linux x8664 + aarch64, Windows x8664, Android. A stable C ABI (speechcorec.h) makes it trivially bindable from Swift / Kotlin / Python — speech-swift is a sibling project that consumes it through a prebuilt XCFramework binary target.
Apache 2.0. Tested on macOS / Linux / Windows in CI; LiteRT runtime is fetched per-platform from the official wheel by a small shell script.
Repo: https://github.com/soniqo/speech-core
Feedback on the interface design / build setup very welcome.
https://redd.it/1tvuifx
@r_cpp
Post #25362
14