Training · learner work

The Call Ended. The Next Step Didn't Have To.

When an AI voice agent finishes a call with a prospect, someone has to decide what happens next. Kes, who joined Colaberry as an intern and is now an AI Systems Architect, built the rules that read what the caller asked for and schedule, cancel or stop the follow-up, and then bounded them when they retried too much.

verified

All student projects
How the routing works, in 101 seconds. Narrated by a synthetic voice; the figures it states are the verified metrics recorded below.
Industry
Education
Capability
Operational AI call intent routing
Status
Shipped
Built by
Kes · Colaberry
Published
2026-09-18

The situation

The system places outbound voice calls for an admissions team through an AI voice agent and analyses each answered call afterwards: a classification, a summary, a consent check. Until 29 Mar 2026 that was where the automation stopped. The analysis job created a task for a person, wrote the summary and updated the lead's state, and what happened next depended on somebody reading the summary and deciding.

The decisions were not alike. A caller who asked to speak to a real person needed a person within hours. A caller whose booking link had failed needed another attempt. A caller who said not to call again needed every scheduled call cancelled and the number flagged. A conversation that ended well under two minutes with no conclusion needed someone to judge whether to try once more or let it go. On the console each of these looked the same: a completed call with a summary attached, waiting.

Rule-based intent detection already existed for the voicemail path, where it read a transcript for phrases like a callback request. Answered calls carried two more signals the voicemail path never had: the platform's own list of actions it had taken during the call, such as an attempted transfer or a failed booking, and the call's duration. The requirement Kes wrote on 29 Mar 2026 was to run the detection after the analysis of every answered call, with those three inputs, and to give each intent one explicit follow-up action instead of a task in a queue.

For a builder the interesting part is the boundary. The analysis step already used a language model and worked; the temptation was to ask the same model what to do next. The requirement drew the line the other way: the model writes the summary, and the decision is patterns, thresholds and an order that a colleague can read in one file and a test can pin. Everything the system does afterwards, and everything this record can measure, follows from that line.

What it had to do

  • After the analysis of every answered call, detect the caller's intent from the transcript, the platform's reported actions and the call's duration, by rule.
  • Give each intent one explicit follow-up action: schedule a call or a message, change the lead's state, or stop the sequence.
  • Decide between competing signals in a fixed priority order, with stop signals ahead of requests.
  • Leave the reason on every scheduled job and the detected intent on every call record, so the dashboard can show both.

What constrained it

  • Detection is rule-based by design; the module header records that no language model is used.
  • Explicit requests from a caller, such as a callback at a stated time, must keep working and are not capped.
  • A conversation the rules cannot interpret must not be treated as a request.
  • The routing runs after an analysis step that depends on an external AI service and a voice platform the team does not control.
The Lead Lifecycle Monitor: each lead's last intent and the next action the rules scheduled
The Lead Lifecycle Monitor: each lead's last intent and the next action the rules scheduledCaptured from the repository's own dashboard (dashboard-ui, the /lead-lifecycle page) running locally at the measured head with empty keys, seeded with thirteen placeholder leads, one per routing outcome. Every name is Demo Contact NN and every number is in the fictional +1555010NNNN range; no real lead appears. The Last Intent and Next Action columns show what the rules wrote for each placeholder call.

Decisions that made the difference

Three choices shaped the system. Each one lives at a specific point in the drawing above.

  1. At Intent detection

    Separate interpretation from action

    The analysis already produced a summary, and the platform already reported what it had tried during the call. Neither was a decision, and letting the summarising model also choose the next step would have made every routing choice unreadable after the fact.

    The model keeps the summary. The decision is a separate module that reads three inputs, the transcript, the platform's executed actions and the duration, through patterns and thresholds, and returns one intent or nothing. The handlers act only on that result.

    Evidence Architecture narrative and the release of 29 Mar 2026: detect_intent with its three inputs, the module header stating no language model is used.

    96%of answered calls since 13 Apr 2026 received a next step by rule, 2,227 of 2,326, and for each one the reason can be read from a pattern table rather than inferred from a model.

  2. At Priority order

    Let priority rules govern the next step

    One call can carry two signals: a callback request and a request not to be called again, or a booking the platform reports as failed and a transcript that reads as uncertain. Something has to win, and a stop has to beat an ask.

    A fixed order, first match wins. The platform's reported transfer or booking comes first; an inaudible transcript next; then the pattern table in its published order: the stop signals (do not call, wrong number, not interested), then enrolled, a request for a person and a lead re-engaging, then the callbacks and a failed booking, then the softer answers, and last the requests to be texted or emailed; partial engagement only when nothing matched and the call was short.

    Evidence The requirement of 29 Mar 2026 lists the order; the detection file at the measured head runs it; the measurement notes count what each route did.

    31 of 34callers who said not to call again were never scheduled again by any route; the three exceptions were campaign calls a day or more later, not intent follow-ups. Whether the rules chose the right intent is not measured: no labelled sample exists, and the record says so.

  3. At Follow-up job

    Bound follow-up and leave a trace

    A handler that schedules another attempt after every inconclusive call has no end, and a mismapped number can absorb attempts for months. By August one contact had accumulated 83 intent retries across five reasons.

    Every scheduled call carries its reason in the payload, every call record its intent, every campaign change an audit row. Partial engagement and inaudible calls mark the lead cold and stop. A failed booking and an unconfirmed transfer retry at most twice. A callback the caller asked for stays uncapped, because a cap would break the legitimate case.

    Evidence Build timeline, 24 Aug 2026; the incident record in the repository; the retry counts in the measurement notes.

    551 to 0retry calls after inconclusive conversations, 551 for 413 contacts before the correction and none since. The trace is what made the audit possible: the 83 retries were found by grouping jobs by their reason.

The build

  1. Repository created
  2. Live-call intent detection and routing shipped
  3. The detected intent stored on every call record
  4. First answered call routed by rule on the live system
  5. Provider actions in dictionary form parsed for booking intent
  6. Low-confidence audio stops retrying
  7. Every remaining uncapped retry loop capped or removed
  8. Measured from the production database, read-only
Notes on 7 of the 8 steps
  • Live-call intent detection and routing shipped. Detection gains the executed actions reported by the platform and the duration of the call; one handler per intent; wired into the analysis job after the summary, where until this commit the job ended in a task for a person, a summary and a lead-state update; the requirement, the architecture note and the runbook; 24 unit tests.
  • The detected intent stored on every call record. A column on the call table and a backfill for earlier calls, so the dashboard can show what the rules decided.
  • First answered call routed by rule on the live system. The earliest call record carrying a detected intent, which is where every production figure on this record starts.
  • Provider actions in dictionary form parsed for booking intent. The platform began reporting executed actions as a dictionary; the parser accepts both forms.
  • Low-confidence audio stops retrying. The first of two fixes from the retry audit: an inaudible call marks the lead cold instead of scheduling another attempt.
  • Every remaining uncapped retry loop capped or removed. Partial engagement no longer retries; failed booking and unconfirmed transfer capped at two; explicit callback requests left uncapped on purpose; 65 tests across the two files; the behaviour of every handler documented in one place.
  • Measured from the production database, read-only. The head of the branch the system runs was cf233ae; the figures on this record were read at 2:26 PM Central that day.
Scheduled Actions: the next seven days of follow-ups the rules produced
Scheduled Actions: the next seven days of follow-ups the rules producedThe /campaign-overview page over the same thirteen placeholder leads: calls and a text scheduled by the handlers, nurture follow-ups with their dates, and the leads the rules closed or flagged. Placeholder rows only; no real lead, number or transcript.

What was built

Interpretation and action are two modules. The analysis job runs first and produces the classification, the summary and the consent check; only then does it hand the transcript, the platform's list of executed actions and the duration to detect_intent. Detection returns one intent and a confidence, or nothing. It is patterns and thresholds in one file, and the file says at the top that no language model is used.

Stack

  • Python
  • Fastapi
  • PostgreSQL
  • Next.js
  • TypeScript
More on what was built

Detection reads its inputs in a fixed order and the first match wins. A transfer or booking the platform reports as executed is trusted before the transcript is read. A transcript under five characters, or only noise, is low-confidence audio. Then the pattern table runs in its published order: the stop signals (do not call, wrong number, not interested), then enrolled, a request for a person and a lead re-engaging, then the callbacks and a failed booking, then the softer answers (interested but not now, call later, uncertain), and last the requests to be texted or emailed. When nothing matched and the call was under two minutes, the result is partial engagement: the rules found no ask in a short conversation.

handle_intent takes the result, cancels the contact's pending jobs, and dispatches to one handler per intent. A handler does one of three things: schedules a launch job for a call, or a job for a text or an email, at a time the rule sets or the caller named; changes the lead's state, to nurture with a date, to cold, to closed, to human transfer, or flags the number; or stops. Every job it schedules carries the intent as its reason in the payload, the call record carries the detected intent, and a campaign change writes an audit row.

The bounds were added on 24 Aug 2026 after an audit found one contact with 83 intent retries across five reasons. Partial engagement and low-confidence audio now mark the lead cold and do not call again. A failed booking and an unconfirmed transfer retry at most twice, then mark the lead cold. A callback the caller asked for is left uncapped on purpose, because a cap would break the legitimate case. The dashboard reads the same tables: the Lead Lifecycle Monitor shows each lead's last intent and next action, and the drill-down shows the jobs, including the one the handler cancelled beside the one it scheduled.

Capabilities

  • Intent detection
  • Rule based routing
  • Follow up scheduling
  • Retry caps
  • Audit trail
  • Operator dashboard

Integrations

  • Voice platform transcript executed actions duration
  • Crm
  • AI analysis service

Data stores

  • PostgreSQL

The measurement

Evidence maturity: measured outcome, read from the production database on 17 Sep 2026 in a read-only session, from 13 Apr 2026 when the first routed call appears. The headline is the correction rather than the release, because it is the one figure with a before and an after of the same measure: 551 retry calls after inconclusive conversations for 413 contacts before 24 Aug 2026, none since. Around it: 2,227 of 2,326 answered calls received a next step by rule; 785 follow-up jobs name the intent that scheduled them; 476 of the 529 that came due completed and 467 of those within fifteen minutes; 31 of 34 callers who asked not to be called were never scheduled again. Two of the eleven candidate figures the build was asked for could not be made: routing correctness, which needs a labelled sample nobody has drawn, and a confirmed transfer, which the voice platform does not report. Both are on the roadmap as not pursued rather than hidden.

Full notes on all 10 metrics
  • 96% of answered calls routed to a next step, 2,227 of 2,326

    Answered calls that received a next step from the rules

    verified

    Baseline
    Before the release of 2026-03-29 the analysis job ended in a task for a person, a summary and a lead-state update; no call was routed by rule. The comparison is with a way of working, not with a number.
    Unit
    share of answered calls
    Sample
    2,326 completed call records from 2026-04-13 to 2026-09-17.
    Methodology
    One read-only query counting completed calls, those with a transcript and those with a detected intent since 2026-04-13, and a second grouping the remainder by whether a transcript exists.

    Limitations

    • A routed call is one where a rule matched; it says nothing about whether the follow-up it chose was appropriate.
    • Calls before 2026-04-13 are excluded because an earlier backfill wrote intents onto old records, which would mix routed calls with relabelled ones.
  • 15 distinct intents across 2,227 routed calls, two thirds of them inconclusive

    What the routed calls asked for, by primary intent

    verified

    Baseline
    None. This sizes the work the rules face rather than comparing it with anything.
    Sample
    2,227 completed calls with a detected intent, 2026-04-13 to 2026-09-17.
    Methodology
    One read-only query grouping completed calls since 2026-04-13 by detected_intent. Intents are mutually exclusive by construction: detect_intent returns the first match in its priority order.

    Limitations

    • A primary intent hides a second one: a lead who asked for a text and a callback is counted once, under whichever rule ranks higher.
    • The share of partial engagement reflects the rule's threshold (under 120 seconds with no other match) as much as the callers.
  • 785 follow-up jobs carry the reason they were scheduled

    Follow-up jobs that name the intent that scheduled them

    verified

    Baseline
    None. This counts a property of the jobs rather than comparing it.
    Sample
    All launch_outbound_call jobs with an intent_reason in their payload, since the release.
    Methodology
    One read-only query grouping launch jobs by payload intent_reason and status; one counting audit_log rows with action campaign_switch (7, from 2026-04-13 to 2026-08-13).

    Limitations

    • Intent reasons name six routes; routes that end in a lead-state change rather than a job (nurture, cold, do not call, transfer) are not represented here.
    • The job carries the reason, not the transcript line that triggered it; tracing further needs the call record.
  • 476 of 529 due follow-ups completed

    Follow-up jobs that ran when they came due

    verified

    Baseline
    None. Before the release no job was scheduled by intent, so there is no earlier rate to compare with.
    Unit
    share of due jobs
    Sample
    785 launch jobs with an intent reason: 476 completed, 53 failed, 255 cancelled, 1 pending and not yet due.
    Methodology
    Read-only queries grouping intent-tagged launch jobs by status, and counting pending ones by whether run_at has passed (0 overdue). Of the 255 cancelled, 11 had a call from the same contact between scheduling and cancellation; the rest were cancelled by a later handler or stop signal.

    Limitations

    • Failed means the worker could not complete the launch; the record does not say why, and the campaign tiers may have called the lead later anyway.
    • Cancellation is a status flip without its own audit row, so the reason a job was cancelled is inferred, not recorded.
  • 467 of 476 completed follow-ups ran within 15 minutes of when they were due

    Completed follow-ups that ran within 15 minutes of their due time

    verified

    Baseline
    None. The delay is a configured policy (two or four hours by route, or the time the lead named); this measures the worker against it, not the policy against an earlier one.
    Unit
    share of completed jobs
    Sample
    476 completed launch jobs with an intent reason.
    Methodology
    One read-only query over completed intent-tagged launch jobs: percentile_cont(0.5) and (0.9) of updated_at minus run_at, and the count under 15 minutes.

    Limitations

    • updated_at is the last status change, which for a completed job is its completion; a job that was touched afterwards would read late.
    • Nine jobs ran more than 15 minutes late; the record does not say why.
  • 266 of 286 requests for a person carried a transfer attempt

    Requests for a person where the voice platform reported a transfer attempt

    verified

    Baseline
    None. This is a distribution across the request cohort, not a comparison with an earlier period.
    Unit
    share of transfer requests
    Sample
    286 completed calls with detected_intent human_transfer_request: 272 before 2026-08-25, 14 from that date.
    Methodology
    Read-only queries over completed calls with detected_intent human_transfer_request, testing the executed_actions text for a transfer and the job table for a transfer_requested job within 24 hours: attempt and follow-up 3; attempt only 263; follow-up only 4; neither 16.

    Limitations

    • A provider-reported attempt is not a confirmed hand-off; the platform payload has no field for the outcome of the transfer.
    • The follow-up rule was capped at 2 per contact on 2026-08-24; the 7 follow-ups span both policies.
  • 0 of 2 contacts booked after a retry

    Contacts whose booking retry led to a booking the platform reported

    verified

    Baseline
    None. The retry route has only ever been taken for 2 contacts.
    Sample
    9 launch jobs with intent_reason booking_retry, for 2 contacts, since the release; 7 calls in all were routed as failed_booking.
    Methodology
    One read-only query joining booking_retry jobs to later calls for the same contact and testing the executed_actions text.

    Limitations

    • A sample of 2 supports no rate; the card is a statement about the trace, not about the route's effectiveness.
    • A booking completed outside the platform (by a person, or through a link) leaves no mark the query can find.
  • 31 of 34 contacts who said do not call were not scheduled again

    Contacts who said do not call and received no further outbound call job

    verified

    Baseline
    None. This is a compliance rate for one route, not a comparison.
    Unit
    share of do-not-call contacts
    Sample
    34 contacts with at least one completed call routed as do_not_call.
    Methodology
    Read-only queries: distinct do-not-call contacts and the existence of a later non-cancelled launch job; those jobs bucketed by gap and reason (all tier calls, all a day or more later); lead_state for the 34 (22 closed and flagged, 11 with no row, 1 unflagged); later call records (4 outbound, 1 inbound).

    Limitations

    • A job is the routing's unit of intent; a call placed by another path would show as a call record without a job, and 4 of the 34 have a later outbound call record.
    • Twelve of the contacts have no lead-state row carrying the flag, so the stop for them rests on the job cancellation alone.
  • 453 of 533 follow-ups were scheduled within ten minutes of the call

    Follow-ups scheduled within ten minutes of the call record

    verified

    Baseline
    None. Before the release the next step waited for a person to read the summary, and no timestamp records when that happened.
    Unit
    share of follow-up jobs
    Sample
    533 intent-tagged launch jobs created within 24 hours of the contact's last completed routed call; 117 more were created a day or more later and are outside the set.
    Methodology
    One read-only query pairing each intent-tagged launch job with the latest completed routed call for the same contact at or before the job, keeping pairs under 24 hours: median 0.1 minutes, 90th percentile 120.3 minutes, 453 under ten minutes. A second query buckets the pairs (453 under a minute, 0 from one to ten minutes, 8 from ten to sixty, 45 from one to three hours, 27 from three to 24) and confirms the one-to-three-hour bucket has a median of 120.4 minutes with none within ten minutes of the paired call's start. A third measures run_call_analysis turnaround (2,322 jobs, median 0.1 minutes, 90th percentile 0.2; 100 over 30 minutes).

    Limitations

    • The pairing is by contact, not by job lineage, so a job created after an unanswered retry is measured from the earlier answered call and reads two hours slow.
    • The 90th percentile of 120.3 minutes is therefore the retry policy showing through the measurement, not the time the rules took to decide.
  • 65 unit tests for live-call intents and handlers, from 24 at release

    Unit tests covering live-call intents and their handlers

    verified

    Baseline
    24 test functions in tests/unit/test_live_call_intents.py at 55ee0fa, the release commit.
    Sample
    The two test files at 9702521.
    Methodology
    git show 9702521:tests/unit/test_live_call_intents.py | grep -c "def test_" (35); the same for tests/unit/test_intent_actions.py (30); and at 55ee0fa for the first file (24).

    Limitations

    • A count of test functions is not coverage; it says how many cases were written, not how much of the routing they reach.
    • This is a build fact, kept out of the headline on purpose.

What happened next

Shipped

  • Rule-based intent detection with provider actions and duration
  • One handler per intent: schedule, change state or stop
  • Detected intent stored on every call record
  • Last intent and next action on the operator dashboard
  • Retry caps and terminal actions after the incident audit

Not pursued

  • A reviewed, labelled sample to measure routing correctness
  • A confirmed-transfer state from the voice platform
Notes on 4 of the 7 items
  • Last intent and next action on the operator dashboard. The Lead Lifecycle Monitor, the contact drill-down and Scheduled Actions read the intent and the job reason; the images on this record are those pages.
  • Retry caps and terminal actions after the incident audit. Correction of 24 Aug 2026, five months after the release: partial engagement and low-confidence audio no longer retry; failed booking and unconfirmed transfer capped at two. The audit that prompted it found one contact holding 83 intent retries across five reasons.
  • A reviewed, labelled sample to measure routing correctness. No labelled sample exists. Route counts and the 65 unit tests say how often a rule fired and that the expected phrases map to the expected intents; neither says how often the rule was right. This record does not claim a correctness figure for that reason.
  • A confirmed-transfer state from the voice platform. The platform reports that a transfer was attempted, not that it connected. Until it does, follow-through for a request for a person can only be measured to the attempt.

Meet the builder

Kes

AI Systems Architect, Colaberry

  1. Intern
  2. Hired by Colaberry
  3. AI Systems Architect

Project contribution

Wrote the requirement and built the detection, the priority order and one handler per intent on 29 Mar 2026, wired them into the analysis job, and stored the detected intent on the call record for the dashboard. Five months later, after an audit found one contact with 83 retries, capped or removed every remaining retry loop and documented each handler's terminal behaviour. The release commits carry no co-author; the later fixes name an AI coding assistant as co-author, which the record states rather than hides.

Skills demonstrated

  • Rule design over model outputKept the language model to the summary and put the decision in inspectable patterns and thresholds; the module header says no language model is used (architecture, 29 Mar 2026).
  • Precedence and stop conditionsA fixed priority order in which the platform's reported actions override the transcript and stop signals rank above requests; the first match wins (architecture).
  • Handler designOne handler per intent that schedules, changes the lead's state or stops, each leaving its reason on the job it created: 785 follow-up jobs carry theirs (measurement).
  • Bounding an automation after an auditTraced 83 retries on one contact to uncapped handlers and gave every route a terminal behaviour on 24 Aug 2026; retries after inconclusive calls went from 551 to none (build timeline, measurement).
  • Testing the rules24 unit tests at release and 65 after the caps, including cap-boundary cases that check below the cap retries and at the cap marks cold (build timeline).

Career facts as confirmed to Colaberry; project contribution from the repository record.

What this project shows

This project shows what an AI Systems Architect is responsible for after the model has spoken: deciding what happens next, in rules a colleague can read, and coming back to bound them when they run too far. Kes's routing gave 2,227 of 2,326 answered calls a next step, and the correction he made five months later took the retries after inconclusive calls from 551 to none. The gap is stated on the record: nobody has yet measured how often the rules chose the right intent.

Build one of these

Start the program that produced this work.

See the program