Field guide · Operations

Run a 30-Day AI Front Office Pilot

A go or no-go decision scored against criteria you committed to before the pilot started, with the evidence attached.

Lessons
9
Read
18 min
Figures
5
Start at lesson 1

By Velaire Health · September 13, 2026 · 18 min read

ShareXLinkedIn

Most pilots in this category end with everyone agreeing it went quite well and nobody able to say what would have counted as failure. That is not a pilot, it is a trial period with a sunk cost attached.

It seems to be goingwell.

day 30, the review

the baseline nobody took

Nobody wrote down what failure would look like, so nothing could fail.

Before you start

  • A measured baseline, ideally from sizing your missed-call gap first
  • Your written call handling rules, including escalation and prohibited responses
  • A vendor willing to run a bounded pilot with an exit
  • One named person with about 90 minutes a week for 5 weeks

What you will end up with

  • A baseline taken during setup week, because after go-live the comparison is gone permanently.
  • Success criteria and kill criteria agreed and dated before anything is switched on.
  • Forty randomly sampled calls reviewed against your written rules, not a vendor dashboard.
  • The handoff timed as carefully as the call, including how many captures needed a clarifying callback.
  • A written decision citing the numbers, and a set of artefacts that make the next evaluation faster.

The default version of this goes badly in a predictable way. A practice switches something on for a month, nobody measures anything beforehand, the vendor's dashboard shows encouraging numbers, and at the end everyone agrees it seems to be working. The subscription continues because stopping would feel like admitting a mistake.

That is not an experiment. An experiment has a thing it is trying to find out, a measurement taken before the intervention, and a stated outcome that would count as failure. A trial without those is a purchase with a delay built in.

Nine lessons below, spread across 5 weeks: 1 week of setup, 4 weeks of running, and a scoring session at the end. The discipline is deliberately front-loaded, because everything that makes a pilot informative has to be decided before it starts.

What a pilot can and cannot tell you#

Thirty days is enough to learn whether a system handles your actual calls correctly, whether your staff can work with what it produces, and whether the handoff functions. Those are the questions worth answering and they are answerable in a month.

It is not enough to establish a revenue effect. Booking cycles, seasonality, and normal variation swamp a 30 day sample, and any vendor presenting a revenue lift from 4 weeks is presenting noise with a confident label on it. Judge operations, not outcomes.

That distinction sets the whole design. Every criterion below is about behaviour you can observe directly rather than a downstream number you would have to infer.

Thirty days answers "does this work correctly". It does not answer "did this make us money". Do not let the pilot be scored on the second one.

The largest confounder in a 30-day window is usually who was working. MGMA polled 357 practices in May 2025 and found front-office roles were the most frequently cited turnover hotspot even among practices whose overall turnover was improving [3]. One resignation inside your pilot window will move every number you are measuring, in the direction that flatters the vendor, and nothing in a dashboard will tell you that is what happened.

01Take the baseline before anything is switched on#

A pilot without a baseline can only be compared against impressions, and impressions are exactly what the pilot exists to replace. This has to happen in the setup week, because after go-live the comparison is gone forever.

  1. Export the last full month of inbound call records and record total volume, missed count, and the missed split between staffed and unstaffed hours.
  2. Record your current time to first human response for an after-hours enquiry, honestly, including the ones that took until the afternoon.
  3. Count how many enquiries in that month produced a booked appointment, and how you know.
  4. Time 10 real calls of the type the pilot will handle, end to end, including the admin afterwards.
  5. Write all 4 numbers on one page and date it.

Step 4 is the one people skip and it is the one that makes the economics real later. A 6 minute call plus 2 minutes of typing is 8 minutes, and knowing that is what turns "it saves time" into a number you can multiply.

Step 2 needs an honest definition before you can record it. Time to first human response means from when the patient first tried to reach you to when a person actually spoke to them, not to when a message was seen. Practices consistently record the second number because it is the one their systems produce, and it can understate the real figure by most of a working day.

Value the time you measured in step 4 while you are there. Published wage data puts the median for medical secretaries and administrative assistants at $22.08 an hour as of 2025 [1], and your own loaded cost is higher. That converts a saved minute into a figure without needing anyone to build a business case.

Time to first response means until a person spoke to them. Not until somebody saw a message.

If you have not done this before, size your missed-call gap is the long version of steps 1 and 3 and is worth the afternoon.

The four baseline numbers
  • Total inbound volume, missed count, and the staffed against unstaffed split
  • Current time to first human response for an after-hours enquiry, honestly
  • Enquiries that produced a booked appointment, and how you know
  • Ten real calls timed end to end, including the admin afterwards
Lesson 1. All four come from records you already have, during setup week.

02Write the success criteria before you start#

The single highest-value 20 minutes in this guide. Criteria written after the fact adjust themselves to whatever happened, and everyone involved will do this unconsciously and in good faith.

Commit to numbers in this shape:

DimensionCriterionHow it is measured
Correctness95% or better of sampled calls handled per your written rulesManual review of sampled calls
Escalation100% of immediate-band calls reached a humanEvery one, checked individually
Capture quality90% or better of captures usable without a clarifying callbackFront-desk marks each one
SpeedTime to first response under a stated targetTimestamp comparison
Staff burdenConfirmation time per capture under 90 secondsTimed sample of 10

The escalation row is the only one at 100% and that is deliberate. A system that routes 97% of urgent calls correctly has failed, because the 3% is the entire reason the band exists.

Sign the page. Not literally, but get the practice owner and whoever runs the front desk to both look at these numbers before go-live and agree they are the test. The agreement is what stops them moving.

Set the numbers at what you actually need rather than at what you hope for. A criterion of 99% correctness sounds rigorous and will fail almost any system including a human front desk, which means the pilot ends in a negotiation about whether the bar was fair. Pick thresholds you would genuinely accept, then hold them.

One criterion deserves a note on measurement. Correctness is scored against your own written rules from the call handling rules guide, not against a general sense of whether the call went well. If you have not written those rules, the correctness criterion has nothing to be measured against and the pilot cannot produce a defensible number.

Where an external comparator exists, borrow it rather than inventing a threshold. A 2022 study of 6 Veterans Affairs medical centres selected for above-average primary care access records the telephone standards those sites were held to: an average speed of answer of 30 seconds or less, and call abandonment under 5% [4]. Numbers you did not choose yourself are much harder to quietly relax in week 3, which is exactly what makes them useful in a success criterion.

03Write the kill criteria as well#

Success criteria tell you when to continue. Kill criteria tell you when to stop early, and a pilot without them will run its full 30 days regardless of what happens in week 1.

  • Any immediate-band call that did not reach a human, ever, at any point
  • Any clinical advice or symptom interpretation given to a caller
  • Any confirmation to a third party that a named individual is a patient
  • Correctness below 80% in the week 2 sample
  • Any patient complaint specifically about the automated handling
  • Front-desk confirmation time above 3 minutes per capture, sustained

The first 3 are absolute and stop the pilot the day they occur. They are not performance issues, they are the failures the whole call handling document exists to prevent, and treating them as data points to be averaged is the wrong response.

Row 4 needs a stated response as well as a threshold. Correctness below 80% in the week 2 sample is not automatically a stop, it is a decision point, and the useful version says who makes that call and by when. Left undefined it becomes a conversation that runs until week 4 arrives and settles it by default.

Row 5 is worth defining narrowly. A complaint about the automated handling is different from a complaint about the practice generally, and conflating the two will either trip the criterion unfairly or let a real signal hide inside normal grumbling. Write down what counts before you need to judge it.

Name who has authority to invoke a kill criterion, and make sure it is not the person who championed the pilot. That is not a comment on anyone's integrity, it is that nobody is a good judge of the thing they advocated for.

The person who can stop the pilot should not be the person who chose the vendor. That is a structural point, not a personal one.

04Pick the narrowest scope that could still fail#

A pilot covering everything tells you nothing, because when it underperforms you cannot say which part. Narrow it until it covers one clearly defined slice that is nonetheless capable of failing.

  1. Pick one call type from your banded list, usually new-patient enquiries.
  2. Pick one time window, usually after close through to opening.
  3. Pick one location if you have several.
  4. Leave everything else exactly as it is, including your existing after-hours arrangement as a fallback.
  5. Write down what is explicitly out of scope, so nobody scores the pilot on it.

Step 4 matters more than it sounds. Running the pilot alongside your existing arrangement rather than instead of it means a failure is an inconvenience rather than an incident, and it means you can stop on the day a kill criterion trips without leaving a gap.

Narrow enough to fail
  1. 1
    One call type
    Usually new-patient enquiries, taken from your banded list.
  2. 2
    One time window
    Usually after close through to opening the next morning.
  3. 3
    One location
    If you have several, the pilot runs at one of them.
  4. 4
    Existing arrangement stays
    Running alongside rather than instead of means a failure is an inconvenience.
Lesson 4. One call type, one window, one location, everything else untouched.

Resist scope creep during the month, however reasonable each request is. Adding a second call type in week 3 means you no longer have 4 comparable weeks, and the person asking will not be the person who has to interpret the result.

Keep a parked list rather than saying no outright. Requests that arrive mid-pilot are usually good ideas arriving at a bad time, and writing them down does two useful things: it stops the conversation without a fight, and it produces a ready-made scope for the second phase if the answer turns out to be yes.

One scoping question is worth settling explicitly before you start, because it will otherwise be settled by whoever is on shift. Decide whether existing patients calling out of hours are in scope or out. Both answers are defensible, the behaviours are quite different, and mixing them means the correctness number is an average across two populations that should have been measured apart.

05Sample real calls, do not read the dashboard#

This is the lesson that most separates a real pilot from a vendor demonstration. A dashboard reports what the system believes happened. Listening reports what happened.

  1. Sample 10 calls a week, chosen randomly rather than by the vendor.
  2. Listen to each in full, or read the complete transcript, not the summary.
  3. Score each against your written rules: correct band, correct capture, nothing prohibited said.
  4. Keep the failures in a list with a one-line note on what went wrong.
  5. Do this in week 1, week 2, week 3, and week 4 without changing the method.

Forty calls over the month is enough to see the pattern and small enough that one person can do it in about 45 minutes a week. Random selection is essential: a sample the vendor picks is a showcase, and a sample of the calls that went wrong is a complaints file. Neither is a measurement.

Read the transcript rather than the summary, always. Summaries are generated by the same system you are evaluating, and a system that misheard a caller will produce a summary that is internally consistent and wrong.

The same principle applies to how you judge an effect size at the end. The evidence literature in this area is a useful corrective: a Cochrane review of appointment reminder trials reported an effect and still concluded the evidence was insufficient to conclusively inform policy [2]. A 30 day sample from 1 practice is much weaker than that, so treat your result as an operational verdict rather than a measured outcome.

Your pilot is not a study. It is a competent operational check, and it should be reported as one.

There is one sampling refinement worth adding in week 3 if you have capacity. Alongside the 10 random calls, deliberately review every call the system itself flagged as low confidence. That is a different question from correctness: it tells you whether the system knows when it does not know, which is the property that makes the escalation rule dependable.

Sampling, four weeks running
  1. Week 1
    Sample 10
    Score against the written rules. Start the failures list.
  2. Week 2
    Sample 10, then review
    Check kill criteria explicitly. Fix configuration only, and date it.
  3. Week 3
    Sample 10
    Change nothing. Resist every reasonable scope request.
  4. Week 4
    Sample 10, then score
    Fill in the criteria page before the vendor's report arrives.
Lesson 5. Ten random calls a week, same method every week, about 45 minutes.

The failures list is the most valuable artefact the pilot produces. It is what you show the vendor, it is what distinguishes a fixable configuration problem from a limit of the product, and it is far more persuasive in a renewal conversation than any aggregate.

06Instrument the handoff, not just the call#

Most of the time saving a practice is buying happens after the call ends. If you only measure the call, you have measured the half that the vendor controls and none of the half your practice lives with.

What to timeWhy
Capture to first human viewDoes anyone actually see it, and when
Human view to confirmationThe real per-item cost of the new process
Captures needing a clarifying callbackThe honest measure of capture quality
Captures actioned same working dayWhether the queue is a queue or a pile

The third row is the one that decides whether this is worth anything. A capture that requires a callback to clarify has replaced one phone call with a different phone call, and if that rate is high the product has moved work rather than removed it.

Have the front desk mark each capture with a single tick for usable or not usable as they process it. It costs nothing, it takes the judgement away from the aggregate, and it produces a number nobody can argue with at the end.

Define usable before the first tick is made, or you will get 4 weeks of inconsistent judgement. A workable definition is that a capture is usable when a person can act on it without contacting the patient for clarification. Not whether it was perfect, not whether the caller was pleasant, just whether it needed another call.

Row 4 deserves an explicit threshold as well as a measurement. Captures actioned the same working day should be near 100% during opening hours, and a rate materially below that says the destination you chose in your handoff design is not being watched. That is a process finding about your practice rather than a verdict on the product, and it is worth separating in the write-up.

07Hold a week 2 review and change almost nothing#

Halfway through, look at the numbers together. The purpose is to catch a kill criterion and to notice a fixable configuration problem, not to start improving the score.

  1. Read the failures list from weeks 1 and 2 out loud.
  2. Check each kill criterion explicitly rather than assuming none tripped.
  3. Separate configuration problems from product limits, and be honest about which is which.
  4. Fix configuration problems, and record the date you changed anything.
  5. Change nothing else.

Step 4's date matters. If you change the configuration in week 2 you no longer have 4 comparable weeks, you have 2 plus 2, and your final score should be computed on the later pair with that stated. Silently changing things mid-pilot and reporting a monthly average is the most common way a pilot flatters itself.

Step 3 is the judgement call. A system saying the wrong practice name is configuration. A system unable to route reliably to a human when uncertain is not, and no amount of prompt adjustment will change it.

08Score it against what you wrote in lesson 2#

At the end of week 4, go back to the criteria page and fill in the actual numbers beside the committed ones. Do this before any discussion, and preferably before the vendor's own end-of-pilot report arrives.

  • Correctness percentage from the 40 sampled calls
  • Every immediate-band call in the month, individually checked
  • Usable-capture percentage from the front desk's ticks
  • Median time to first response, against your baseline
  • Median confirmation time from the timed sample
  • Any kill criterion that tripped, with the date
  • Total captures, and how many are still unactioned

The last row catches a specific failure that otherwise hides. A pilot that captured 90 requests and left 30 unactioned did not save the practice time, it created a backlog, and the backlog will not appear in any correctness measure.

Reading the vendor's report afterwards rather than first is a small discipline with a large effect. Their numbers are usually accurate and they measure what their system did, which is a subset of what you care about.

Where a pilot flatters itself
Keeps it honest
  • Random sample, chosen by you
  • Full transcripts rather than summaries
  • Criteria agreed and dated before go-live
  • A sceptic doing the scoring
Quietly invalidates it
  • Calls the vendor selected
  • Reading the dashboard instead of listening
  • Criteria written after seeing the results
  • Configuration changed mid-month and averaged across
Four habits that turn a measurement back into an impression.

09Decide, and write down why#

The decision is a sentence, and it should reference the criteria rather than a feeling. Write it down whichever way it goes, because the reasoning is what you will want in 6 months when the question comes back.

  • The decision, in one sentence
  • Which criteria were met and which were not, with the numbers
  • What you would need to see to change the decision
  • What you learned about your own front office, separate from the vendor
  • The date, and who made the call

The fourth row is the one that has value regardless of outcome. Most practices running this properly learn something about themselves that has nothing to do with the vendor: that captures sit for a day before anyone looks, that nobody was tracking escalations, that the recall list is the real problem.

The third row is the one that keeps a no from being permanent. "We would revisit this if the usable-capture rate reached 90%" is a specific, checkable condition that lets a vendor come back with something real, and it stops the decision hardening into a general belief that this category does not work.

Be careful how the decision gets communicated internally, particularly after a no. Staff who were sceptical will hear a no as vindication and staff who were hopeful will hear it as a verdict on their judgement. The numbers on the page are the antidote to both readings, which is another reason to circulate the scored criteria rather than the conclusion alone.

A no is a good outcome, cleanly reached. It cost you 5 weeks and about 8 hours of one person's time, and you now have a baseline, a failures list, and a criteria page that make the next evaluation much faster.

What a no still buys you
5 weeks
elapsed, including setup
about 8 hours of one person's time
40 calls
sampled and scored against your rules
the failures list is reusable evidence
6 months
the baseline stays valid
criteria page reusable on the next vendor verbatim
The artefacts outlive the decision, which is why a clean no is cheap.

What to hand the vendor at the end#

Give them the failures list and the scored criteria page, whichever way you decided. A vendor receiving 40 sampled calls with annotated failures is getting something genuinely useful, and how they respond to it is itself information about who you would be working with.

Ask them one question when you hand it over: which of these failures is configuration and which is a limit of the product today. A vendor who answers that honestly, including naming the ones they cannot fix, is worth more than one who promises all 40 will be resolved by next month.

If the answer is yes, keep the measurement running at a lower cadence. Five sampled calls a month rather than 10 a week is enough to notice drift, and drift is real: a configuration that was correct in March meets new treatments, new staff, and new call reasons by September.

If the answer is no, keep the artefacts. The baseline stays valid for about 6 months, the call handling rules stay valid indefinitely, and the criteria page can be reused verbatim on the next vendor. That is the compounding benefit of running one of these properly, and it is why the call handling rules come first.

ShareXLinkedIn
Good questions. Clear answers.

Questions about this

Why 30 days rather than 90?

Thirty days is long enough to see whether calls are handled correctly and whether staff can work with the output, which are the answerable questions. Ninety days mostly adds cost and sunk-cost pressure without adding clarity, because a revenue effect is not reliably measurable at either length against normal variation.

The vendor's dashboard already reports all of this. Why sample manually?

Because a dashboard reports what the system believes happened, and the failure you most need to catch is the system being confidently wrong. A misheard caller produces a summary that is internally consistent and incorrect. Listening to 10 calls a week is the only method that catches that class of error.

Should we tell patients they are part of a pilot?

Disclosure of automated handling is a separate question from disclosure of a pilot, and the first one matters more. Decide your identity and disclosure rule as part of your call handling document, and check your state law, since at least one state has specific requirements for communications about clinical information.

What if the vendor will not agree to a bounded pilot?

That is useful information delivered early. A vendor confident in their product on your call types has little to lose from 30 days with a defined exit. Reluctance usually signals either an onboarding cost they do not want to absorb twice or a product that needs longer to look good, and both are worth knowing.

Our front desk is against this. Does that invalidate the pilot?

It biases the softer measures, so design around it. Correctness scored against written rules and escalation checked individually are objective and hard to skew. Have a sceptic do the sampling rather than an advocate, since a sceptic who reports 95% correctness has produced a far more credible number.

Keep reading

More from the blog

Start here

What is an AI front office?

The plain-language explainer: what it does, what it deliberately doesn't, and how it compares with voicemail, an answering service, a phone tree, and hiring.

Read the guide
Let’s make room for better care

Your next caller
could be your
next patient.

See how Velaire would handle the calls your practice misses. A personal walkthrough, built around your questions.

Let’s talk about your practice
Illustrative warm, quiet practice reception at the end of the day