Metaviz AI
AI Automation SaaS

Three agents, three lines, and a rule that the AI only ever talks

AI Phone Receptionist for a Tax Resolution Practice

A voice receptionist for a one person tax practice. Three agents answer three campaign lines, recognise the caller before the phone is picked up, work out how long is left on an IRS notice, and book a real calendar slot during the call.

Overview

A single enrolled agent running a tax resolution practice takes more calls than one person can answer. Every missed call on a paid campaign number is money spent on a click that never became a consultation, and every answered call that turns out to be the wrong fit is billable time gone.

We built a receptionist that answers three separate campaign lines. The number dialled decides which agent picks up, so nobody gets bounced between agents and asked the same questions twice. Before the call connects, a Python service looks the caller up against a synced copy of the CRM and hands the agent five facts, so it knows on its first word whether it is speaking to a client or a stranger.

During the call the agent checks how long is left on an IRS notice, reads real slots off a live calendar, and sends the intake text. After the call a pipeline writes 82 structured fields into the CRM, pages the principal when a case is genuinely urgent, and puts everything else into one evening digest.

The rule the whole build rests on is that the AI only talks. It does not calculate, quote from memory, or improvise, because a wrong date in tax work can cost somebody their case. Every spoken price sits in version control, and a test fails the build if a single word of it changes.

- Relevant keywords -
  • Voice AI
  • Conversational AI
  • Retell AI
  • FastAPI
  • Python
  • Make.com
  • Twilio
  • Calendly API
  • CRM Integration
  • Deadline Triage
  • Appointment Booking
  • Test Driven Delivery
Challenge & Solution

Solving real problems with smart engineering

The caller hears ringing while the system is still thinking

Problem

The telephony platform fires its pre-call webhook before the call is answered, and the caller listens to silence until it returns. A live CRM round trip measured 789 ms at the median and 1,440 ms at worst, against a contractual budget of one second.

Solution

The lookup moved into a purpose built FastAPI service backed by a contact store loaded into memory at startup. Warm lookups are effectively instant against a 26 ms cold start. The live CRM backend stayed behind the same interface with a 400 ms timeout and the local store as fallback, and the choice was settled by a benchmark rather than a preference.

An agent that quotes the wrong price is worse than no agent

Problem

Six fee constants, four qualification floors and 151 quoted lines across three prompts, all edited in a vendor dashboard where a change cannot be diffed and nobody can say when it happened.

Solution

All three prompts live in version control and are pushed to the platform by a sync tool that updates by id and never creates. A test parses every quoted line out of the prompts and checks it against the constants module, then reads the published agent back rather than the draft, because a push can report success while callers still hear the old script.

The model must never do date arithmetic

Problem

Ten IRS notice types, each with its own statutory clock. Thirty days of appeal rights on one, ninety days to petition on another, twenty one days on a bank levy. The clock decides the fee quoted and whether the caller reaches a human. A language model asked to count days will answer confidently and be wrong.

Solution

A triage module owns a documented precedence ladder and computes the deadline, the days remaining and the priority in plain Python against offline test vectors. The agent supplies only what the caller reads off the letter. Three real faults were found this way, including an agent that accepted a notice date in the future and invented a deadline from it.

Five vendor platforms, all able to fail quietly

Problem

A contact sync reported success on every run for two weeks and had never written a single row, because its request was returning 401 and the automation platform does not treat that as an error. A booking guard meant to refuse appointments past a caller deadline was dead for weeks for a similar reason.

Solution

Standing checks now exercise the live path instead of reading configuration: the lookup, the client path, both numbers, one caller one record, the stored ids, the hourly sync, and a change in the CRM arriving in the store. Automation blueprints are snapshotted and committed before every change.

One caller turning into two records

Problem

Six of nine captured callbacks were a different number from the line the person rang in on, and the callback was overwriting the original. The same caller came back unrecognised, with their name on one record and their real number on a nameless second one.

Solution

Both numbers go on the record now, with the originating line pinned in a field nothing later overwrites, and deduplication matches on either. The client marker moved out of CRM group membership and into an explicit field, because every caller lands in that group and membership as a marker would have let any second time caller walk past qualification.

Nothing may be spoken that was not returned

Problem

Numbers get recycled, partners share handsets, and caller ID can be spoofed. A summary of a previous call read out to the wrong person is a confidentiality breach in a practice handling enforcement matters.

Solution

The pre-call response is deliberately small: five variables, each a string, returned on every call whatever the service finds. Detail from earlier calls is written to the CRM for the principal and never returned to a live agent, and a test asserts the exact set of variables, so putting one back has to be a deliberate act.

How it works

What happens on a call

The same path runs on every line, which is what makes the record at the end predictable.

Step 1

Before the ring

A FastAPI service resolves the caller against the synced CRM copy and returns five facts.

Step 2

The lane answers

The number dialled decides which of the three agents picks up. There is no router.

Step 3

Qualification

The caller is measured against the revenue and balance floors for that lane.

Step 4

Deadline triage

The notice type and date go to a Python module that returns days remaining and priority.

Step 5

Booking

Live slots are filtered to the booking window, and no slot past the caller deadline can be offered.

Step 6

Intake text

One SMS goes out, bound to the offer that preceded it.

Step 7

The record

82 fields are written to the CRM on both the create and the update path.

Step 8

Alerting

Critical callers page the principal. Everything else lands in the evening digest.

Features

What it does

The call

  • Three campaign numbers, three agents, no router and no handoff
  • Caller resolved against the CRM before the call connects
  • Qualification against hard revenue and balance floors, with the release final
  • Deadline triage across ten IRS notice types, each with its own clock
  • Live calendar search filtered to the booking window before a time is read aloud
  • Booking guard that refuses any slot past the caller deadline
  • Warm transfer to the principal only inside the agreed window

The record

  • 82 structured fields written to the CRM on both the create and the update path
  • Urgent callers paged within five minutes, everything else in one evening digest
  • Originating line and callback number both stored, with the original pinned
  • Every quoted line held in version control and checked by an automated test
  • No detail from a previous call is ever returned into a live one
  • CRM changes reach the lookup store in about fifteen seconds
Tech Stack

Built with

Backend

  • Python 3.12
  • FastAPI
  • Uvicorn
  • Pydantic v2
  • httpx
  • SQLite contact store

Voice and telephony

  • Retell AI
  • Twilio SIP trunking
  • Programmable SMS
  • A2P 10DLC registration

Automation

  • Make.com
  • Folk CRM REST API and webhooks
  • Calendly API
  • SMTP alerting

Quality

  • pytest
  • mypy strict
  • Ruff
  • GitHub Actions CI
  • uv
Results & Value

What actually changed

The practice answers every call on every line, quotes only what it has authorised, and knows which callers are running out of time before anyone picks up the phone.

  • Beat a contractual latency gate by two orders of magnitude, settled by benchmark rather than preference
  • Made it impossible for the agent to quote a price the practice had not authorised
  • Took post-call capture from partial to complete by root-causing three separate silent failures
  • Shipped the verification as a deliverable: an acceptance suite, live path checks and blueprint snapshots
01Headline outcome
26 msPre-call lookup, coldAgainst a contractual one second budget. Warm lookups are effectively instant.
02
447
Automated testsWith strict typecheck and lint clean alongside them.
03
87
Acceptance scenarios31 of them graded blocker severity.
04
82
Fields captured per callUp from 9, 18 and 21 across the three lanes.
05
10
IRS notice types triagedCritical, expedited and standard, computed in code rather than guessed.
06
151
Spoken lines under testEvery quoted line across the three agent prompts.
07
63%
Fewer operations per callFourteen operations on a standard call, down from thirty eight.

Have a similar project in mind?

Let's talk about how AI can transform your service business.

Talk to us about a system like this