End-of-turn detection in production
Five 2026 stacks compared on the same recorded audio set. How do LiveKit, Pipecat, Vapi, Retell, and a VAD-only baseline handle slow speakers, backchannels, and noise?
手法論文、独自のベンチマーク、比較分析、フィールドレポート——現在も継続的に公開中。すべての研究に、データセット、スクリプト、事前登録済みの研究手法を添付しています。ゲート付きのPDFも、ホワイトペーパーを装った営業資料もありません。
各研究は6つのパイプライン段階を経て進みます。カードは研究の進行に合わせて更新されます。
Five 2026 stacks compared on the same recorded audio set. How do LiveKit, Pipecat, Vapi, Retell, and a VAD-only baseline handle slow speakers, backchannels, and noise?
P50 / P90 / P99 end-to-end "user stops speaking to first audio byte arrives" latency across four stacks, controlled for region, cold start, and prompt budget.
Where Spanish-English voice agents drop the turn. Four leading bilingual STT stacks scored on a 100-utterance gold-labelled set with controlled code-switch points.
01
敗者を明確にする。
当社の計算機がXを3倍上回り、Yと互角なら、それを率直に明記します。
02
データを公開する。
生データ、スクリプト、すべての図表を再現できるワンライナーを含むリポジトリ。
03
利害関係を開示する。
ベンチマーク対象のツールを再販売したり提携したりしている場合は、最初にその旨を明言します。
04
限界を正直に述べる。
"さらなる研究が必要"という定型句は使いません。論文から導き出せない具体的な事項を列挙します。
私たちが研究する分野のための、情報密度の高い1ページリファレンス。番号付きのトピックブロック、本当に重要な数式、本番環境で実際に起きる障害パターン。
Ten dense topic blocks covering the production voice-agent stack in 2026: pipeline shape, end-of-turn detection, backchannel handling, latency budgets, streaming TTS, code-switching, mid-turn tool calls, audio + token cost models, failover, and the eval harness that catches regressions before customers do.
Ten compact blocks on the production RAG stack: embeddings, chunking, hybrid retrieval, reranking, citation grounding, caching, index sizing math, per-query cost, the eval harness, and the failure modes that show up after launch.
Ten dense blocks on the 2026 browser-automation fingerprint surface: detection vectors, canvas + WebGL signals, audio fingerprint, TLS JA3/JA4, behavioural signals, the stealth library stack, residential proxy economics, the CAPTCHA solver pipeline, total cost per successful session, and the eval harness used to track block rates over time.