Buyer's guide
When Not to Automate Your Phone
Four situations where an AI front office is the wrong purchase, what the cheaper fix is in each, and the one shape of practice where the arithmetic genuinely works.
A written decision between five options, scored against your own after-hours volume and your own urgency rules, with a dated review point.
There are five things a practice can do about the hours nobody is there, and one of them is nothing. This is how to choose between them using your own numbers rather than a vendor's framing of the question.
a vendor called on Tuesday
somebody saw the invoice on Friday
back to voicemail, again
The after-hours question usually gets answered by whoever asked it most recently. A vendor calls and the answer becomes automation. A partner complains about a lost patient and the answer becomes an answering service. Somebody works out what an answering service costs and the answer goes back to voicemail.
None of those is a decision. They are responses to whoever spoke last, and the reason they cycle is that the practice never established what it actually needs the cover to do.
Seven lessons, about 3 hours, ending in a written decision with a review date. The output may well be that you should do nothing, and the guide is written so that outcome is available rather than embarrassing.
There are five real options, and they differ along only two axes that matter: whether a person is involved at the moment the call happens, and what the practice receives afterwards. Everything else about them, including price and including how they are marketed to you, follows from where they sit on those two.
Doing nothing means the phone rings out. Voicemail records a message. An on-call rota routes to a person who carries a phone. An answering service has somebody else answer and send you a message. Automated cover answers, captures a structured request, and escalates what it should not handle.
It is worth noticing that four of the five end the same way: with something arriving that a person at the practice has to act on the next morning. Only an on-call rota resolves anything at the time, and even that resolves only the calls the person on call can actually deal with from wherever they are.
That shared ending is why the handoff, rather than the answering, is where most of these options succeed or fail. A practice comparing them purely on what happens at 8pm is comparing the smaller half.
Ranking those by cost is easy and mostly beside the point. Ranking them by what they do for the specific calls you actually receive after hours is the exercise, and it requires knowing what those calls are.
Five options, and one of them is nothing. A guide that cannot reach that conclusion is a sales document.
Everything downstream depends on this and most practices have never looked at it separately. An after-hours call and a daytime call are different populations and averaging them hides the entire question.
Use a full month rather than a fortnight. After-hours volume is lumpy in a way daytime volume is not, and a single quiet or busy week will move the count enough to change which option looks sensible.
Step 4 is the one that changes the answer most and almost nobody does it. A caller who rang at 8pm and got through at 9:20 the next morning was not lost, and a practice where most after-hours callers do that has a much smaller problem than its raw count suggests.
All 5 steps come out of the same export, so do them in one sitting rather than returning to it.
Step 5 is worth doing even if you expect the answer to be low, because a very low voicemail count against a much higher missed-call count is itself informative. It means your after-hours callers are choosing to hang up rather than to leave a message, which tells you something about how much patience they have and therefore about how likely a callback is to reach them.
Step 2's distinct-caller count deserves as much attention as the total. Twenty after-hours calls from 18 different people and 20 from 7 people are completely different situations: the first is a coverage gap affecting 18 potential patients, and the second is a handful of persistent callers who will almost certainly reach you eventually. Vendors quote the first number and the second is the one that decides whether you have a problem.
Step 3 usually shows the volume is concentrated rather than spread. A common shape is most of the week's after-hours calls landing in the 2 hours after close, with very little overnight, and that shape changes what a sensible option looks like.
If you have not done this before, the fuller version of it is sizing your missed-call gap, and it is worth the afternoon before you spend anything.
Volume alone does not tell you what the calls will tolerate. A 2024 Talkdesk survey of 1,000 US adults put comfort with AI scheduling routine appointments at 42%, and the share of patients with sensitive issues who would rather book through a chatbot than speak to someone at 67% [2]. It was paid for by a company selling contact centre software, so read it as a direction. The direction is that the after-hours mix in an elective practice is more automatable than the daytime mix, not less.
This is a clinical decision rather than an operational one and it belongs to whoever holds clinical responsibility. It also determines which options are available to you, so it comes before pricing rather than after.
| Category | Must reach a human | Acceptable delay |
|---|---|---|
| Clinical concern or possible complication | Yes, always | Minutes |
| Post-procedure question | Yes | Same evening |
| Booking or rescheduling | No | Next working day |
| General enquiry or price question | No | Next working day |
| Anything the caller describes as urgent | Yes, regardless of category | Minutes |
Get this signed off rather than merely discussed. The categories in this table are the ones any option will be configured against, and if they were agreed informally in a corridor then the person configuring the system is making clinical judgements on the practice's behalf without knowing it.
The last row is not a hedge, it is the rule that keeps the others safe. Categories are assigned by whoever built the system and urgency is assessed by the person calling, and when those two disagree the caller wins.
Assign categories yourself, and let the caller assign urgency. When those two disagree, the caller wins.
Once this table exists, some options remove themselves. If your first row genuinely requires a human within minutes, then voicemail and doing nothing are not candidates regardless of what they cost, and the choice narrows to three.
Keep the table short. A practice that produces 11 categories has built something nobody at a desk will apply at 8pm on a Friday, and the value of this lesson is entirely in whether the person or system covering your phone can actually hold it in mind while a caller is talking.
That narrowing is the point of doing this before looking at prices. Deciding what you need and then pricing it produces a different answer from pricing things and then deciding what you need.
There is a published yardstick for the speed half of this. A 2022 study of 6 Veterans Affairs medical centres chosen for above-average primary care access records the standards those sites were held to: an average speed of answer of 30 seconds or less, and call abandonment under 5% [3]. Nobody expects that at 11pm. It is useful anyway, because it forces you to write down the number you do expect instead of leaving it as "quickly".
Now cost them, using your own volume rather than list prices, and include the costs that do not appear on an invoice.
| Option | Direct cost | The cost nobody quotes |
|---|---|---|
| Nothing | None | Whatever share of after-hours callers do not come back |
| Voicemail | Near zero | Staff time listening and returning, and an unowned queue |
| On-call rota | Overtime or an allowance | Goodwill, and the cost of it failing during leave |
| Answering service | Per call or per minute | Morning callback work, since a message is still a task |
| Automated cover | Monthly fee | Confirmation time per capture, and somebody to own the queue |
The first row's cost is the hardest to establish and the most argued about, which is why lesson 1 spends its effort there. It is not the count of after-hours calls; it is the share of distinct after-hours callers who never reached you at all, multiplied by a conversion rate you have discounted and a patient value from your own records. Any of those four inputs guessed generously turns a modest number into an alarming one.
The right-hand column is where comparisons usually go wrong. Voicemail looks free and is not, because a person listens to every message and rings back; an answering service looks cheap per call and produces the same callback work the next morning.
Voicemail looks free because its cost is paid in minutes rather than on an invoice. Nobody bills you for the 45 minutes.
Value the staff time honestly while you are here. Median pay for medical secretaries and administrative assistants runs $22.08 an hour as of 2025 [1], and your loaded cost is higher, so 45 minutes a morning of callbacks is a real recurring number.
| What it looks like | What it also costs | |
|---|---|---|
| Voicemail | Free | Listening and callbacks, into an unowned queue |
| Answering service | Cheap per call | The same callbacks, next morning |
| On-call rota | An allowance | Goodwill, and failure during leave |
| Automated cover | A monthly fee | Confirmation time, and an owner for the queue |
Do the arithmetic per option at your own volume rather than comparing headline prices, because volume changes the ranking. An answering service billed per call is cheap at 20 after-hours calls a month and stops being cheap at 200, while a fixed monthly fee behaves the other way round. The crossover point is specific to you and is usually not where either vendor's example sits.
The on-call row deserves particular care because its cost is mostly not financial. A rota that depends on a person's willingness works until that willingness runs out, and it tends to run out without notice.
With the requirements from lesson 2 and the costs from lesson 3, score each option rather than arguing about it. The scoring is what stops the decision reverting to whoever speaks last.
The fourth line matters for the review in lesson 7 rather than for the decision now. An option that records what it did gives you something to check in 6 months; one that does not leaves the review as a discussion about impressions, and impressions are exactly what this whole guide exists to replace.
The seventh line is worth more than it looks. Reversibility is a real feature, and options differ enormously on it: a rota can be stopped on Friday, and a 12 month contract cannot.
The second line catches the option practices most want to choose. An on-call rota scores well on almost every other row and fails this one, and the failure is not hypothetical: it arrives every time somebody takes leave, and it arrives permanently the day the person carrying the phone decides they have had enough.
The sixth line is easy to skip and occasionally decisive. An option that only operates from 9pm is worth very little to a practice whose after-hours volume lands between 5:30pm and 7pm, and that mismatch is invisible unless you did lesson 1 step 3.
The third line is the one that separates options that look similar. An answering service and automated cover both mean the caller speaks to something at 8pm, and one produces a paragraph while the other produces fields, which decides how much work lands on your desk in the morning.
The fifth line is where enthusiasm meets arithmetic. Run your volume at a discounted conversion rate rather than an optimistic one, because an option that only pays under the best case is an option you are buying on hope.
Every option fails. Choosing one means choosing which failure you would rather have, and that is a much more useful question than which one performs best when everything works.
The fifth line is the one practices find uncomfortable to look at. Doing nothing does not produce a bad report, an angry patient, or a complaint, because the person who could not reach you simply went elsewhere and never told you. It is the only option on this list whose failures generate no evidence at all, which is why it feels safer than it is.
The last line is the important one. Four of the five options end with something arriving for a person to action, and if that queue has no owner then the option you chose barely matters, for the reasons in voicemail is a queue nobody owns.
There is a second question worth asking about each failure, which is whether it fails loudly enough for you to find out in a week or quietly enough to persist for a quarter. A loud failure is annoying and self-correcting. A quiet one accumulates, and by the time somebody notices, the practice has months of it behind them and no record of what was lost.
Think about which failure your practice would actually notice. Voicemail's failure mode is invisible, automated capture's failure mode is visible and occasionally embarrassing, and a lot of practices would rather have the visible one because it can be corrected.
Every option fails. The choice is which failure you would rather have, and which one you would actually notice.
Whatever you choose, deploy it to the smallest slice that could still fail, and keep whatever you have now running alongside it. This costs almost nothing and removes most of the risk from the decision.
Step 3's instruction not to change anything mid-way is the one people break with good intentions. A week 2 tweak that improves things leaves you with 2 weeks of one arrangement and 2 weeks of another, and a monthly average describing neither. If something is badly wrong in week 2, stop the trial rather than adjusting it.
Step 2 is what makes a bad choice cheap. If the new arrangement fails in week 2, the practice is inconvenienced rather than exposed, and you can stop on the day rather than at the end of a contract.
Step 1 has a subtlety. Choose the window where your volume actually is rather than the window that is easiest to cover, because a trial run in a quiet window will tell you the option handles quiet windows and nothing more. If most of your after-hours calls land in the 2 hours after close, that is the trial, even though it is also the hardest slot.
Step 4 is the discipline that most distinguishes a real evaluation from an impression, and the full version of it, with pre-committed success and kill criteria, is running a 30 day pilot.
The last lesson is what stops this cycling back in 6 months when somebody new asks the question. A decision that exists only as a purchase is not a decision anyone can revisit intelligently.
The fifth line is the one that makes the review honest. Writing down in advance what failure looks like means the review is a comparison rather than a discussion about whether everyone feels it is working.
The fourth line is what stops the review being a conversation about feelings. "We expected this to convert about 6 after-hours enquiries a month into booked requests" is checkable in a way that "we expected it to help" is not, and writing it down at the point of purchase costs a sentence.
Name who owns the review as well as when it happens. A review date with no owner is a date that passes, and the question comes back 3 months later through whoever happens to raise it, which is precisely the cycle this guide was written to stop.
Six months is the right interval because the inputs move. Opening hours change, volume moves seasonally, a competitor opens or closes, and an answer that was correct in September can be wrong by March without anybody doing anything wrong.
Whatever you choose, write down what the caller will be told about it. Pew Research Center, polling 3,488 US adults in June 2026, found 56% wanted to be told when AI was used to schedule their appointment, and 63% wanted more say in whether it was used at all [4]. Scheduling is the most tolerated use of AI there is, and it is still not one most people want discovering by accident.
For a real share of practices, particularly single-location ones with light evening volume and persistent callers, the correct answer is to leave things as they are. That is a legitimate conclusion and this guide is written to make it reachable.
Nothing is also the only option that is free to reverse, which is worth weighing alongside its risks. A practice that decides to do nothing and writes down why has lost 3 hours and gained a documented position; a practice that signs a 12 month contract on the same evidence has committed real money to a question it had not answered.
The shape looks like this: a small after-hours count, most of it concentrated in the first hour after close, a high proportion of those callers reaching you the following morning anyway, and a low case that does not clear any option's cost. The reasons that pattern makes buying a mistake are worked through in when not to automate your phone.
It is worth saying explicitly that this can change, and usually changes in one direction. Practices grow, evenings get busier, and a decision that was correct at 20 after-hours calls a month becomes wrong at 60 without anyone noticing the crossing. That is the entire reason for the review date rather than for a permanent conclusion.
If that is your shape, the useful output of these 3 hours is a written note saying so, with the numbers attached and a review date on it. The next time somebody raises the question, you have an answer rather than an argument, and you will know exactly which number would have to change for the answer to be different.
The plain-language explainer: what it does, what it deliberately doesn't, and how it compares with voicemail, an answering service, a phone tree, and hiring.
See how Velaire would handle the calls your practice misses. A personal walkthrough, built around your questions.
Let’s talk about your practice