End-of-turn detection in production
Five 2026 stacks compared on the same recorded audio set. How do LiveKit, Pipecat, Vapi, Retell, and a VAD-only baseline handle slow speakers, backchannels, and noise?
方法論論文、原創基準測試、比較分析與實戰報告,持續產出中。每篇發表的研究都完整附上資料集、腳本與預先註冊的方法論。沒有設門檻的 PDF,也沒有披著白皮書外衣的業務簡報。
每項研究皆經歷 6 個流程階段。卡片會隨著研究進度即時更新。
Five 2026 stacks compared on the same recorded audio set. How do LiveKit, Pipecat, Vapi, Retell, and a VAD-only baseline handle slow speakers, backchannels, and noise?
P50 / P90 / P99 end-to-end "user stops speaking to first audio byte arrives" latency across four stacks, controlled for region, cold start, and prompt budget.
Where Spanish-English voice agents drop the turn. Four leading bilingual STT stacks scored on a 100-utterance gold-labelled set with controlled code-switch points.
01
我們點名輸家。
如果我們的計算機比 X 快 3 倍且與 Y 持平,我們會明確指出。
02
我們公開數據。
存放原始數據、腳本及一鍵重現所有圖表的指令碼的程式碼庫。
03
我們揭露利益衝突。
若我們轉售或與被評估的工具合作,我們會事先說明。
04
限制陳述誠實無華。
沒有「需要更多研究」的套話。清楚列出該論文無法得出的具體結論。
針對我們研究的各個領域,提供高密度的一頁式參考資料。編號主題區塊、關鍵公式,以及實際環境中會出現的失敗模式。
Ten dense topic blocks covering the production voice-agent stack in 2026: pipeline shape, end-of-turn detection, backchannel handling, latency budgets, streaming TTS, code-switching, mid-turn tool calls, audio + token cost models, failover, and the eval harness that catches regressions before customers do.
Ten compact blocks on the production RAG stack: embeddings, chunking, hybrid retrieval, reranking, citation grounding, caching, index sizing math, per-query cost, the eval harness, and the failure modes that show up after launch.
Ten dense blocks on the 2026 browser-automation fingerprint surface: detection vectors, canvas + WebGL signals, audio fingerprint, TLS JA3/JA4, behavioural signals, the stealth library stack, residential proxy economics, the CAPTCHA solver pipeline, total cost per successful session, and the eval harness used to track block rates over time.