WORKING DEMO / VOICE

AGENT RUN APPS / OUR OWN WORKING DEMO

Did the voice agent actually say it?

We built an AI patient simulator to test a voice conversation, record both speakers, and check the result against the actual audio. Four rehearsals exposed missing replies, early turns and a goodbye that existed only in a tool call.

RECORDED AUDIO01:12 / REHEARSAL 04

An appointment at a fictional clinic.

Both speakers are AI voices. All patient and office details are synthetic. This is a local demo, not a client deployment; no telephone number was dialed.

Recorded September 8, 2026 · 72.20 seconds.
Adam’s listening review of voice and pacing is pending.

WHAT WE CHECKED

A tool call cannot prove that words were spoken.

In an early run, the patient’s finish request claimed it had said goodbye. Its recorded speech had not. Comparing generated text with the recording caught a failure that the tool result alone would have missed.

The local Python runner exchanges audio with OpenAI’s audio services, saves each speaker on a separate channel, and transcribes those recorded channels. Playback events help explain what happened. The audio remains available for a person to review.

  1. Set the scenarioOne appointment request and fixed fictional facts.
  2. Record both speakersKeep the original audio and per-run events.
  3. Compare the evidenceCheck the recording, transcript and expected behavior.
  4. Change and repeatMake a specific fix, then record another rehearsal.

FOUR RECORDED REHEARSALS

What failed, and what changed.

  1. 01

    The conversation ended before the words did.

    The receptionist hit its output limit and lost parts of two replies. The patient called the finish tool, but the recording contained no final confirmation or goodbye.

    We increased the output allowance, asked for shorter replies, and checked that closing audio had finished playing before ending the session.

  2. 02

    A promised readback was still missing.

    The next recording included a repaired goodbye, but the patient said it would repeat the details without actually doing so. It also started replying during two receptionist turns.

    We increased the silence allowance from 650 to 1,000 milliseconds and required the closing to state the agreed details aloud.

  3. 03

    The details were right. The speaker’s role was wrong.

    The third rehearsal included the confirmation and goodbye, but the patient used the receptionist’s phrasing: “You’re confirmed.”

    We made the closing instruction explicit: stay in the patient role and confirm the appointment in first person.

  4. 04

    The fourth recording contains the complete closing.

    The patient repeats the agreed time, address and parking information, then says goodbye. The playback log confirms that the final audio played.

    One closing repair was still needed to supply the missing speech. We kept that limitation in the review; voice, pacing and the final transition still await Adam’s listening approval.

REQUIREMENTS REVIEW / SEPTEMBER 8

Review each requirement against its evidence.

This is how we review a build: identify the requirement, locate its evidence, and name what is still untested. Test counts describe implementation checks, not customer results.

Hold a two-sided conversation

Rehearsed locally

Four local rehearsals exchanged actual audio between an AI patient and a fictional AI receptionist.

Keep evidence of what was heard

Recording retained

Both speakers were recorded on separate channels. The displayed transcript comes from those recorded channels; generated text was retained separately as provisional.

Confirm the details before finishing

Observed with repair

The latest recorded closing repeats the agreed details in first person and says goodbye. A single repair supplied the missing closing and waited for playback.

Check the implementation and setup

Checked offline

The offline suite passed 274 tests, all 38 dependency pins matched, and local package build/install checks passed. These checks use mocked transport and do not prove telephone behavior.

Exercise a range of scenarios

Coverage incomplete

Twelve scenarios with 48 proposed checks are prepared. Only fictional appointment scheduling has recorded audio evidence so far.

Review voice quality and telephone operation

Pending

Adam’s voice and pacing review is pending. Twilio connectivity, public callbacks and behavior against an external voice agent remain unverified.

THE RECORDED WORDS

Read alongside the audio.

Open the recording transcript

Automatically transcribed from the two recorded channels; human review is pending. Segment times are approximate and some ranges overlap. They do not measure turn latency or prove simultaneous speech. Every name, date and location below is fictional demo data.

  1. AI receptionist0.00–12.00s

    Maple Grove Demo Clinic. This is Casey. How can I help?

  2. AI patient0.00–21.68s

    Hi, I'd like to make an appointment for a general consultation next week.

  3. AI receptionist12.00–18.00s

    Please share your full name and date of birth together so I can set up the fictional booking.

  4. AI receptionist18.00–27.00s

    Then I'll offer a time for next week.

  5. AI patient21.68–35.68s

    My name is Avery Linden and my date of birth is April 16th, 1987.

  6. AI receptionist27.00–42.00s

    I can offer Tuesday, September 15, 2026 at 2.30pm. Does that work for you?

  7. AI patient35.68–37.68s

    Yes, that works for me.

  8. AI patient37.68–54.68s

    Could you please share the clinic address and parking information?

  9. AI receptionist42.00–48.00s

    Your appointment is confirmed for Tuesday, September 15, 2026 at 2.30pm.

  10. AI receptionist48.00–53.00s

    The clinic is at 100 Example Lane and free parking is behind the building.

  11. AI patient54.68–59.68s

    Okay, I'll repeat the details back and wrap up the call.

  12. AI patient59.68–71.68s

    I will be there Tuesday, September 15th, 2026 at 2.30 p.m. at 100 Example Lane and I'll use the free parking behind the building. Thank you. Goodbye.

Download the transcript

TEST YOUR OWN WORKFLOW

Bring the conversation that needs to work.

We can scope the scenarios, expected behavior, evidence and human handoff before building or connecting a voice workflow.

Describe your workflow