Skip to main content

Interruptible Conversational Voice AI on the Edge: Build, Deploy, and Measured Results

Usage notice

This is an open-source reference implementation, not a certified product. Speech accuracy has been measured on RK3576 and on Orin NX (Qwen3-ASR int4) only.

What this solution does​

Someone walks up to a device and talks to it. The device answers out loud; if they start speaking again halfway through the answer, it stops immediately instead of finishing the sentence.

Built for places where a person speaks to a machine with their hands busy: service desks, exhibits and kiosks, robot voice front ends, smart-home and room terminals.

  • Speech stays on the device

    Recognition (Qwen3-ASR) and synthesis (Matcha-TTS) run on the edge host. With a fully local preset the conversation model runs there too, and the device works offline after the first start.

  • Plug in the mic and interrupt

    The reSpeaker XVF3800 does echo cancellation in hardware and is detected automatically over USB, including hot-plug. Speak while the answer is playing and the device stops.

  • Open source, works with your backend

    The code is open. The conversation goes through an OpenAI-compatible API, so pointing it at your own knowledge base, agent or ordering backend is one URL change; the speech layer stays as is.

  • Measured in Chinese and English

    Measured on RK3576: 1.05% character error rate on short Chinese utterances, synthesis at 0.204× realtime. RK3588 and Jetson hosts offer 30 languages.

What hardware you need​

Four things on site: a microphone array, a speaker and a voice host; if conversation text must not leave the site, also an accelerator card or a host large enough for a 4B model.

① The microphone array must do acoustic echo cancellation (AEC) in hardware.

MicrophoneNotes
reSpeaker XVF3800reSpeaker XVF3800 USB 4-Mic Array
USB array with hardware AEC, noise suppression and beamforming
The validated default. Both the 2-channel and 6-channel firmware layouts are recognized and the processed channel is selected automatically. Other arrays fall back to capture channel 1 and need an acoustic test on site before going live.

② A speaker: USB or 3.5 mm, on the same device. Do not mute the microphone during playback to avoid echo; that also disables barge-in.

③ The voice host runs recognition, synthesis and the resident agent, and determines which languages you can offer:

Voice hostLanguages it can serveWhen to pick it
reComputer RK3576reComputer RK3576
Rockchip NPU, Qwen3-ASR W8A8 + Matcha
Chinese, EnglishLowest-cost host for the full local speech stack; the measured data under "Performance and measured data" comes from this board
reComputer RK3588reComputer RK3588
Rockchip NPU, adds Kokoro RKNN for TTS
All 30You need the 28 languages beyond Chinese and English, or plan to add an RK1828 card for local conversation later
reComputer J3011reComputer J3011 (Orin Nano 8GB)
Qwen3-ASR int4 + Matcha on the GPU
All 30The host also runs other AI workloads and needs GPU headroom. Do not run a local 4B model on it
reComputer J4012reComputer J4012 (Orin NX 16GB)
Speech plus Qwen3.5-4B on the same host
All 30Conversation text must not leave the site

reComputer R2000 series also runs a CPU speech stack (sherpa-onnx, English only); it is not offered in the configurator on the reference design page yet. It does not support Chinese, and ASR and TTS share four CPU cores.

④ Fully local conversation on the RK route: add an RK1828 / RM182X PCIe NPU card to an RK3588 host to run Qwen3-4B. The card needs its own 12 V supply, and the host needs the driver and device node. Only one large model can be resident on the card at a time.

Also: the first start needs internet access and free disk (at least 25 GB on Orin NX), and the deployment tool must be able to reach the host over the network.

How to deploy on site​

Install the microphone and speaker first, then run the deployment wizard.

1. Placing the microphone and speaker​

Check that the microphone has hardware echo cancellation

Barge-in, turn detection and keeping the microphone open during playback all require a capture channel with the speaker's output already removed in hardware. A microphone without hardware AEC causes false interruptions or an echo loop, and no software setting fixes that.

  • Use the reSpeaker XVF3800 by default. If you use another array, run an acoustic test on site first; unknown arrays default to channel 1.
  • Run acceptance with the speaker at normal room volume. A terminal that passes at low volume can fail at operating volume for acoustic reasons unrelated to the models.
  • The reSpeaker can be plugged in before deployment or hot-plugged after the agent is up. The agent selects the capture device by USB product identity, ignores HDMI/DP pseudo-inputs, and recovers from unplug/replug without a container restart.

2. Installing the software​

The SenseCraft Solution app deploys to the host over SSH (or locally, if you are on a Jetson Orin with JetPack 6.2). Per-device steps are in the deployment guide; the outline is four steps.


  1. Choose a preset: cloud / OpenAI-compatible, or fully local. With a cloud preset, conversation text goes to an external endpoint; with a fully local preset, nothing leaves the site.
  2. Choose the conversation language: the deployment resolves (language, device) to one speech profile before starting the other services. An unsupported pair, such as Chinese on reComputer R2000 series, exits with code 2 and stops the whole docker compose up; nothing starts halfway.
  3. Fill in the endpoint and persona: base URL, key and model ID for the cloud preset (defaults: the Beijing-region Qwen endpoint with qwen3.5-flash), plus the system prompt. You can switch from Always listening to Wake word required and type any short Chinese or English phrase; the open-vocabulary sherpa-onnx detector in the image compiles it locally at startup.
  4. Verify in the dashboard: the web dashboard on port 18000 shows listening / thinking / speaking / barged-in.

Acceptance: three turns in the real room at real speaker volume, interrupting 0.5–1 s after each answer starts. Check that the old reply stops immediately and the interrupting utterance is not lost.

Time: about 30 minutes for a cloud preset. Fully local presets take longer because the first start downloads the model files; after one successful online start, images and model files are cached and the device runs offline.

Available interfaces​

The agent calls a streaming OpenAI-compatible Chat Completions API. Point LLM_BASE_URL at your own service to integrate: a RAG endpoint over your documents, an agent framework with tool calls, a robot command layer, an ordering or ticketing backend. The speech layer does not change.

  • Hosted model: keep the defaults or replace base URL, key and model ID. The endpoint must support streaming Chat Completions.
  • Your own service: implement the same interface. The agent sends the transcript as the user turn and streams the reply into synthesis, so the first sentence starts playing before your service finishes generating.
  • Speech layer only: to build a different agent on top, call the duplex WebSocket and offline endpoints below directly.

Complete port and endpoint list​

All services use host networking, so <host> is the voice host's own address.

EndpointWhich deploymentWhat it carries
ws://<host>:8621/v2v/streamevery presetThe duplex session: PCM in, transcript and TTS PCM out, plus the abort a barge-in triggers
POST http://<host>:8621/asrevery presetOffline whole-file transcription, no VAD and no streaming. The offline accuracy figures under "Performance and measured data" are measured here
POST http://<host>:8621/ttsevery presetSynthesis; the x-rtf response header carries the realtime factor
GET http://<host>:8621/healthevery presetReadiness; used as the Compose healthcheck
http://<host>:18000every presetWeb dashboard: turn state and the transcript of each turn
http://<host>:1828/v1, /healthRK3588 + RK1828 local presetOpenAI-compatible Chat Completions for the on-device Qwen3-4B
http://<host>:8000/v1, /healthOrin NX local presetOpenAI-compatible Chat Completions for the on-device Qwen3.5-4B
LLM_BASE_URL (outbound)cloud presetAny OpenAI-compatible endpoint

Both local routes expose the same interface as the cloud route, so switching changes only LLM_BASE_URL. Unless it points outward, there is no broker and no cloud component in the data path.

Performance and measured data​


Runtimes and key parameters​

The duplex protocol and the Agent are the same on every host; only the backends that run ASR and TTS differ.

Speech hostASR backend / modelTTS backend / modelResolved profile
RK3576rk.asr — Qwen3-ASR, RKNN encoder + RKLLM decoder, W8A8rk.tts — matcha-icefall-zh-en, ORT acoustic + RKNN Vocosrk3576-default
RK3588 (zh / en)rk.asr — samerk.tts — matcha-icefall-zh-enrk3588-default
RK3588 (other 28)rk.asr — samerk.tts — Kokoro v1.0 hybrid, RKNN INT8 decoder front + CPU ONNX prefix/tailrk3588-kokoro-rknn
Orin Nano 8GB / Orin NX 16GB (zh / en)jetson.trt_edge_llm — Qwen3-ASR 0.6B, int4jetson.matcha_trt — matcha-icefall-zh-en, bf16/fp16 Vocosjetson-edgellm-v091-matcha
Orin NX 16GB (other 28)jetson.trt_edge_llm — Qwen3-ASR 0.6B, int4jetson.trt_edge_llm — Qwen3-TTS CustomVoice, int4jetson-edgellm-v091-customvoice
reComputer R2000 series (English)cpu.sherpa_asr — sherpa-onnx streaming zh-en, int8 CPUcpu.sherpa — sherpa-onnx CPU voice, int8rpi5-default

The fully local Orin NX preset uses the TensorRT-Edge-LLM v0.9.1 image; engines built for other TensorRT / JetPack versions are not loaded. On the RK side, model files are pulled at first start.

Language × device matrix. Deployment picks a profile from (language, device). Of the 18 cells, 15 are deployable and 3 are unsupported; of the deployable cells, 5 have end-to-end measurements and in the other 10 each component has on-device measurements. The unsupported cells: Chinese on reComputer R2000 series, and the "other 28" languages on RK3576 and reComputer R2000 series. Chinese does not fall back to Whisper: on every board measured, Whisper's Chinese CER is 35–56%. RK3576 has only the Matcha zh-en TTS voice, so it can transcribe the other 28 languages but cannot synthesize an answer.

Three parameters that affect deployment results:

  • barge_in_min_speaking_ms 500, barge_in_min_chars 2: a barge-in counts only once the reply has started playing and the interruption is longer than one syllable.
  • playback_drain_enabled true: the Agent stays in SPEAKING until the local playback buffer is empty. With it off, speech during the tail of the reply is treated as a new turn and the old reply keeps playing.
  • Server-side end of turn: silero, 400 ms of silence + QWEN3_ASR_FRONTEND_EOU_MIN_AUDIO_S=2.5, ending the turn at the first natural pause; a long sentence with a pause in the middle gets an answer to its first clause only.

reComputer RK3576 series measurements​

Conditions: Debian 12 + Rockchip BSP, profile rk3576-default; the test script and the speech container run on the same device and talk over 127.0.0.1:8621, so the numbers exclude network transfer.

ASR on the offline endpoint (POST /asr, whole file, no VAD, no streaming):

LanguageSetnError rate
Chineseshort5CER 1.05%
Chineselong 10–20 s5CER 9.62%
Englishshort5WER 3.65% / CER 1.11%
Englishlong 10–20 s5WER 6.99% / CER 4.16%

Every English long file came back as full multi-clause text; the remaining errors are recognition deviations ("3:2" transcribed as "three to two"), not truncation.

ASR in a live session (/v2v/stream, factory defaults ASR_MAX_NEW_TOKENS=64, ASR_FINAL_STOP_ON_PUNCT=1), scored against the full reference text:

LanguageSetnError rate
Chineseshort5CER 29.09%
Chineselong 10–20 s5CER 84.06%
Englishshort5WER 16.95% / CER 10.16%
Englishlong 10–20 s5WER 63.38% / CER 62.58%

The high live-session error rate comes from the server-side end-of-turn decision, not decoder truncation: with ASR_MAX_NEW_TOKENS raised to 256 and ASR_FINAL_STOP_ON_PUNCT set to 0, the transcripts in both languages are byte-for-byte identical.

TTS (POST /tts, read from the server's x-rtf response header, Matcha ORT acoustic + RKNN Vocos, icefall zh-en voice):

LanguagenReal-time factorRange
Chinese50.2040.190–0.216
English50.1940.158–0.216

English audio per sentence is 2.5–4.3 s long.

Live-session latency (echo mode, no LLM in the chain):

MetricLanguagenValue
End of speech → first TTS frame (stop_to_tts_audio)Chinese41575 / 1624 / 1596 / 1580 ms (about 1.6 s)
ASR final → first TTS frame (final_to_tts_audio)English5p50 1127 ms, mean 1126 ms, 1021–1232 ms
End of speech → ASR final (eos_to_final)English5mean 2837 ms

The measured Chinese end of speech → ASR final (stop_to_final) value arrives later than the spoken reply and is not cited.

Reproduce: docs/perf/rk3576-matrix-20260906.md and docs/known-issues/rk3576-v2v-multi-utterance-timeout.md in the openvoicestream repository; English eos_to_final with asr_stream_ws_bench.py.

reComputer J40 series measurements​

MetricValueConditions
Chinese ASRCER 0 on the golden set, streaming and offlineQwen3-ASR 0.6B int4

Known degradation​

  • Microphone without hardware AEC. Every number above was taken on a hardware-AEC capture channel. Without it the microphone causes false barge-ins or feedback, and no configuration value makes up for it.
  • Pauses mid-sentence. On the same audio, the offline endpoint gives CER 9.62% and the live session 84.06%; the gap comes from the turn-taking policy, and content after a mid-sentence pause gets no answer.
  • Acceptance volume below working volume. A terminal that passes at low volume can fail at working volume for acoustic reasons.

Next steps​

  • Tune the server-side end-of-turn detection so that a mid-sentence pause no longer ends the turn early.
  • Measure end-to-end latency and accuracy for the other 10 deployable cells of the language × device matrix.

Data and asset sources​

  • FLEURS: the ASR corpus for the RK3576 measurements, CC BY 4.0, 5 short and 5 long clips per language, sha256-verified. TTS used 5 self-written sentences per language. The source for the cross-device Whisper comparison is docs/perf/whisper-cross-device-20260827.md in the openvoicestream repository.
  • Raw run records: docs/perf/rk3576-matrix-20260906.md in the openvoicestream repository; every RK3576 figure above maps to a record there.
  • Speech models: Qwen3-ASR, Matcha-TTS (icefall zh-en voice), Kokoro v1.0, Qwen3-TTS CustomVoice and sherpa-onnx each keep their upstream licence terms. Check the licence of each model you actually deploy before commercial shipment.
  • The corpora are not distributed with the repository; obtain them separately.
Loading Comments...