Turn-taking for real speech

Let people finish.

Speech has pauses, repetitions, and restarts. StutterTurn is an engineering prototype for comparing when a voice agent decides to answer — examining the space between a pause and a finished turn.

A pause can still be part of a turn.

Illustrative sequence 00:06.00concept loop
Concept illustration. No recorded timing or audio is represented here.

Inspect what
we have.

The replay-v2 handoff checks the recording-to-event path. It is an engineering check, with human review and model evaluation still ahead.

Recorded-source replay v2

2026-10-10 · Replay v2 supersedes the first replay
Saved logs executed on bcf50d1
Sources
5 SEP-28k recordings
Runs
10 replay logs across stock and patient
Setting changed
Silence gate: 0.2 s → 1.0 s
Observed
Stock decided “complete” inside the labelled gap on 4 of 5 clips; 3 of those replies were cancelled before any audio
Candidate model
None exists yet; the console lists one only when the backend supplies it

Labels have not been human-reviewed. These runs do not establish improved interruption rates or response delay. Audio receipt is logged; speaker playback is unverified.

Inspect the pinned replay evidence

Same recording. A different wait.

The console puts the caller and the agent on one timeline, so a turn decision and reported response audio can be inspected separately.

  1. 01

    Choose a recorded source

    Replay a SEP-28k clip — a real recording of a person who stutters, identified on screen. Keep that source fixed when you change configurations.

  2. 02

    Compare stock and patient

    Start a fresh run for each. Both use the same model; the patient setting waits longer through silence.

  3. 03

    Follow the event record

    Inspect speech, holding, turn decisions, and reported audio receipt together. A decision alone does not mean the caller heard a reply.

Watch the turn.
Inspect the record.

The judge view brings the comparison, console, and evidence into one place.

Open judge view