Training · learner work

Recovering the Calls the Workflow Lost

A voice call could finish while the event saying so never arrived. Kes, who joined Colaberry as an intern and is now an AI Systems Architect, traced the silent failure, recognised that a one-off script was not a system, and built detection, operator controls and scheduled replay in five days. Production has since measured what the system recovers, and where it still falls short.

verified

All student projects
How the recovery works, in 94 seconds. Narrated by a synthetic voice; the figures it states are the verified metrics recorded below.
Industry
Education
Capability
Operational AI voice call event recovery
Status
Shipped
Built by
Kes · Colaberry
Published
2026-09-18

The situation

The system places outbound voice calls for an admissions team and learns what each call produced from a completion event sent back by the voice platform. Everything downstream, from the lead's next step to the record a human sees, depends on that event arriving.

For the admissions representative the failure did not look like a failure. The console showed that a call had been launched and nothing after it: not whether the prospect answered, reached voicemail or never picked up, and not what the lead should get next. A finished call that needed follow-up and a call that simply had not ended yet looked identical, and the only way to tell them apart was for an engineer to go into the voice platform's own records by hand. Each of those calls was a decision waiting on a fact the pipeline had lost.

In late April 2026 Kes's own specification records a burst: hundreds of calls fired in the same second, and a share of the completion events for finished calls never arrived. The calls had happened. The pipeline did not know. The specification is the builder's account of what happened; the production database gives the count. Over the three days of the incident, 533 of 1,644 launches, one in three, had no completion event.

The first response was a script: an engineer exported the affected contacts to a CSV and replayed them by hand, twice, on consecutive days. It repaired 245 of the 533; 288 never got a result. It worked, it did not scale, and it left no trace an operator could see the next morning. Recovery still depended on someone noticing and someone intervening.

What it had to do

  • Detect a completed launch with no completion event within a defined window.
  • Let an operator recover a call in one action, and record that it was done.
  • Recover automatically on a schedule, bounded per cycle, without duplicating work.

What constrained it

  • The completion event comes from an external voice platform the team does not control.
  • A recovered result had to go through the same processing as a live one, or two code paths would drift.
  • Replaying a call that had in fact arrived must not create a second record or a second downstream action.
  • Operators, not engineers, needed to be able to see and act on the gap.
The webhook delivery panel with three missing completions and their four recovery actions
The webhook delivery panel with three missing completions and their four recovery actionsCaptured from the repository's own dashboard (dashboard-ui, the /queue page) running locally at the pinned commit, seeded with synthetic rows: twelve launches, nine with a completion event, three without. Every identifier shown is a placeholder (demo-contact, Demo Campaign); no real lead appears. The figures on the panel are what the system computed from those rows, not typed in.

Decisions that made the difference

Three choices shaped the system. Each one lives at a specific point in the drawing above.

  1. At Detection query

    Make missing events visible

    A finished call that needed follow-up and a call that had simply not ended yet looked identical on the console. Only an engineer checking the voice platform by hand could tell them apart.

    Detection became a query, not a new event stream: a completed launch with no call event inside the window is a failure. The same query feeds the delivery-drop alert, the operator panel and the delivery percentage on the dashboard card.

    Evidence The failures query, the alert, the panel and the dashboard percentage landed on 28 April (build timeline); the panel and the card are the two screenshots on this record.

    28 AprQuery, alert, panel and dashboard card shipped the day after the diagnosis. Operators, not engineers, see the gap the next morning, and every lost event since 30 April has been counted by that definition: 604 of 14,510 launches.

  2. At Same-pipeline replay

    Reuse the existing pipeline

    A second code path for recovered results would drift from the live one, and replaying an event that had in fact arrived must not create a second record.

    The original call record is fetched read-only, its status mapped with the worker's own normaliser, and the ordinary processing job scheduled with a recovery tag. The unique dedupe key already on the call-event table makes a replay of a delivered event reuse the existing row.

    Evidence Architecture narrative; three of the 19 recovery test cases cover exactly this and predate the feature, which reused the guarantee rather than adding one.

    0duplicate call records across 339 recovered calls. The one place the two paths can still drift, campaign inference inside the recovery path, is named on the record.

  3. At Recovery audit row

    Measure the result and its limits

    A repair that leaves no trace cannot be judged, and the first script left none an operator could see.

    Every recovery, advance and ignore writes an audit row. Every outcome figure on this record is read from those rows with the system's own definition of a missing event, and the limits are stated beside it.

    Evidence Measurement notes: 586 of 604 resolved (301 recovered, 285 advanced as no-answer, 18 unresolved); 289 of the 301 recoveries followed the enhancement of 10 September; 4 duplicate CRM tasks from early operator replays.

    97%of flagged events resolved, with the limits beside the figure: resolved is not the same as recovered, most recoveries belong to the later release, and the 18 open cases and 4 duplicate CRM tasks stay in view.

The build

  1. Repository created
  2. Incident observed; call scheduling spread into slots
  3. Detection: the failures query, the delivery-drop alert, the panel, and the dashboard percentage
  4. Manual bulk recovery by script, two incidents, two scripts
  5. Operator recovery: fetch the source call and replay it; VM Left and No Answer actions
  6. Automatic recovery job with sixteen unit tests; Ignore action; campaign inference
  7. Architecture, data model and runbook updated
Notes on 3 of the 7 steps
  • Incident observed; call scheduling spread into slots. The root cause, a burst of simultaneous calls, was mitigated first. Recovery for the events already lost came next.
  • Manual bulk recovery by script, two incidents, two scripts. An engineer replayed affected contacts from a CSV. It worked and did not scale.
  • Operator recovery: fetch the source call and replay it; VM Left and No Answer actions. Three fixes the same day: the response envelope, the status vocabulary, and re-fire suppression once a case is resolved.

What was built

Detection is a query, not a new event stream: a completed launch job whose contact has no call event in the window is a failure, and a ten-minute bucket delivering under eighty percent raises an alert. The same query feeds the operator panel and the delivery percentage on the dashboard card.

Stack

  • Python
  • Fastapi
  • PostgreSQL
  • Next.js
  • TypeScript
More on what was built

Recovery reads the original call record from the voice platform through a read-only client, maps its status into the pipeline's own vocabulary with the worker's normaliser, and schedules the ordinary processing job with the recovered payload, tagged as a recovery. The pipeline that handles a live event handles the replay; there is no second one.

Two guards keep replay safe. The call-event table carries a unique dedupe key, so a replay of an event that did arrive reuses the existing row. And every recovery, advance or ignore writes an audit row that the automatic job and the alert both read, so the same case is not recovered twice and an alert does not fire again for two hours after a resolution.

The scheduled job runs every five minutes over the last three hours, recovers at most ten cases per cycle, defers calls still in progress, and treats a call the platform cannot find as a no-answer to advance. Campaign inference inside the recovery path is a local mirror of the live normaliser rather than a shared function; that is the one place the two paths can drift, and it is listed under limitations.

Capabilities

  • Event recovery
  • Webhook monitoring
  • Idempotent replay
  • Operator controls
  • Scheduled jobs
  • Audit trail

Integrations

  • Voice platform API
  • Crm

Data stores

  • PostgreSQL

The measurement

Evidence maturity: measured outcome. The figures are read from the production recovery audit table on 2026-09-16 with the system's own definition of a missing event. Since the first recovery action shipped on 2026-04-30: 14,510 calls launched; 604 (4.2%) lost their completion event; 586 resolved, of which 301 were recovered and replayed, 285 advanced as no-answer, and 18 remain unresolved. 289 of the 301 recoveries followed the stuck-call enhancement of 2026-09-10, so the figure belongs to the system as it runs today. 96% of recoveries were made by the scheduled job, in a median of 34 minutes; no recovered call produced a second call record, though four early operator replays produced a second CRM task; 91% of automatically recovered calls went on to a further pipeline step. During the incident that started this work, the same definition finds 533 lost events in 1,644 launches, of which the engineer's script repaired 245.

Full notes on all 9 metrics
  • 97% of lost completion events resolved, from 46% by hand

    Lost completion events resolved by the system

    verified

    Baseline
    46% (245 of 533) during the incident of 2026-04-27 to 29, repaired by an engineer running a script from a CSV on two consecutive days. 288 launches from those three days never got a result.
    Unit
    share of detected missing events
    Sample
    14,510 completed launches from 7:00 PM CDT on 2026-04-29 to 8:14 PM CDT on 2026-09-15, of which 604 (4.2%) had no organically delivered completion event.
    Methodology
    Mirror of the recovery job's detection query, run read-only against production. Recovered = manual_webhook_recovery row for the contact; advanced = manual_advance; ignored = manual_webhook_ignore for the job; unresolved = none of these. Replayed events are excluded from "delivered organically" by their source tag.

    Limitations

    • The detection window is the system's, not an independent one: a completion event that arrived more than four hours late counts as missing.
    • 289 of the 301 recoveries happened after 2026-09-10, when stuck-call recovery shipped (a later release than the pinned commit); before it, most missing events were resolved as no-answer.
    • Contacts launched more than once can attach a later resolution to an earlier launch.
  • 4.2% of launches lost their completion event, from 32% during the incident

    Launches whose completion event never arrived on its own

    verified

    Baseline
    32% (533 of 1,644) across the incident of 2026-04-27 to 29, by the same definition.
    Unit
    share of completed launches
    Sample
    14,510 completed launches, 2026-04-30 to 2026-09-16.
    Methodology
    The detection query of the recovery job, applied to every completed launch in the window rather than only the recent ones it looks at in production.

    Limitations

    • The rate is not a delivery rate for the voice platform: it counts by the detection window and includes calls that were never placed.
    • The incident figure is a three-day burst; the after figure is four and a half months, so the comparison is between a bad week and ordinary operation.
  • 96% of recoveries automatic (334 of 347)

    Recoveries performed by the scheduled job rather than a person

    verified

    Baseline
    No before-state to compare: this sizes the split between the two recovery paths, and before the system there was one path, a script run by an engineer.
    Unit
    share of recoveries
    Sample
    347 recovery rows and 457 no-answer advance rows, 2026-04-30 to 2026-09-16.
    Methodology
    select action, operator_id, count(*) from audit_log where action in (manual_webhook_recovery, manual_advance) and created_at >= 2026-04-30 group by 1, 2.

    Limitations

    • The operator replayed some calls more than once in the first two days (one call seven times), so rows overstate operator recoveries of distinct calls: 13 rows, 5 calls.
  • median 34 minutes, p90 47 minutes

    Time from launch completion to automatic recovery

    verified

    Baseline
    The incident repairs ran on 2026-04-29 for calls made on 2026-04-27 and 28: between one and two days.
    Sample
    295 launches recovered by the scheduled job, 7 by the operator.
    Methodology
    percentile_cont(0.5) and (0.9) over extract(epoch from min(audit.created_at) - launch.updated_at)/60, grouped by operator.

    Limitations

    • Includes the job's built-in twenty-minute wait and its five-minute cycle.
    • Multi-launch contacts inflate the tail; the maximum is meaningless and is not shown.
  • 0 duplicate call records in 339 recovered calls; 4 duplicate CRM tasks, all from operator replays

    Recovered calls that produced a duplicate downstream record

    verified

    Baseline
    No before-state to compare: the dedupe guard existed before recovery did. This checks that recovery never defeated it.
    Sample
    339 distinct recovered call ids; 349 recovery rows.
    Methodology
    For each distinct platform call id on a recovery row: count(*) from call_events where call_id = it; count(*) from scheduled_jobs where job_type = create_crm_task and entity_id = it and status = completed.

    Limitations

    • The CRM-task duplication sits one step past the guard: the guard prevents a second call record, and the operator path could still schedule a second task off the same call before it was closed.
  • 91% continued (304 of 334)

    Automatically recovered calls whose pipeline went on to act

    verified

    Baseline
    No before-state to compare: before the system, an unrecovered call had no next step at all, which is the point of the headline figure.
    Unit
    share of automatic recoveries
    Sample
    334 automatic recoveries, 2026-05-01 to 2026-09-16.
    Methodology
    exists(select 1 from scheduled_jobs where entity_type = call and entity_id = call_id and job_type <> process_call_event and status = completed and created_at >= recovery.created_at - 1 minute).

    Limitations

    • 30 recovered calls had no further call-level step; the recovered statuses (254 voicemail, 59 completed, 12 no-answer, 6 busy, 3 failed) do not all warrant one.
  • 5 recovery paths, from none in the product

    Ways a lost completion event can be recovered

    verified

    Baseline
    None in the product. Before 2026-04-30 a lost event was found by hand and repaired by an engineer running a script from a CSV, on two consecutive days.
    Sample
    The repository at the pinned commit; the four panel actions and the scheduled job registration.
    Methodology
    Count the action controls rendered by the failures panel component and the recovery job registered at worker startup. The before-state is dated by the first recovery script commit against the first operator-action commit.

    Limitations

    • This counts what exists, not what was used. The repository keeps no record of how many events each path recovered.
    • The automatic job caps itself at ten recoveries per cycle and defers calls still in progress, so "automatic" does not mean "every case".
  • 19 test cases in 4 files

    Test cases that exercise the recovery behaviour

    verified

    Baseline
    No comparison intended: this sizes the tested surface at the pinned commit.
    Sample
    tests/ at the pinned commit, 52 files.
    Methodology
    grep -c "def test_" on the recovery job test file gives 16; the three dedupe-on-replay cases are named individually in three other files. Everything else under tests/ is counted only to state the denominator.

    Limitations

    • A test count is not coverage. Sixteen of these test one file; the operator-facing half of the system has none.
    • Three of the nineteen predate recovery by seven weeks: the dedupe guarantee was already there and recovery reused it.
  • 53 commits over 5 days

    The recovery system was built in five days

    verified

    Baseline
    No comparison intended: this dates the work, 2026-04-27 to 2026-05-01.
    Sample
    Commits from 2026-04-27 to 2026-05-01 inclusive, filtered to the scoped paths.
    Methodology
    git log --format=%H --since=2026-04-27 --until=2026-05-02 -- <35 scoped paths> | wc -l. Most commits carry an AI coding assistant as co-author, which is stated on the record rather than hidden.

    Limitations

    • The window is chosen by the story, not by the repository; commits after 2026-05-01 that touched these paths are excluded on purpose.

What happened next

Shipped

  • Detection query, delivery-drop alert and dashboard percentage
  • Operator recovery panel with four actions and an audit trail
  • Automatic recovery job, bounded per cycle, with unit tests
  • Stuck-call recovery: a second generation for calls that reported in progress and never finished
  • Recovery counts, from the audit table

Not pursued

  • Tests for the operator recovery path
  • Shared campaign inference between the live and recovery paths
  • Independence from the voice platform
Notes on 5 of the 8 items
  • Stuck-call recovery: a second generation for calls that reported in progress and never finished. Shipped on the main branch on 2026-09-10, after the pinned commit, and running in production: 289 of the 301 recoveries measured on this record came through it. Not part of the build this record describes; named here because the figures include it.
  • Tests for the operator recovery path. The manual recovery function, the read-only call fetch, the recovery endpoints and the panel have no tests at the pinned commit. The automatic cycle and the dedupe guarantee do.
  • Shared campaign inference between the live and recovery paths. Recovery mirrors the live normaliser's campaign inference locally rather than calling it. Status normalisation is shared; this one field is not.
  • Recovery counts, from the audit table. The job still does not persist its per-cycle counters, but every recovery, advance and ignore is an audit row, and those rows are what the outcome figures on this record are read from.
  • Independence from the voice platform. Recovery depends on the platform's call API being reachable and truthful. A call the platform cannot find is treated as a no-answer.

Meet the builder

Kes

AI Systems Architect, Colaberry

  1. Intern
  2. Hired by Colaberry
  3. AI Systems Architect

Project contribution

Created the project and built the detection, recovery and replay architecture: the failures query and delivery-drop alert, the operator panel, the read-only fetch and replay through the existing pipeline, and the scheduled recovery job. The record states, rather than hides, that most commits carry an AI coding assistant as co-author.

Skills demonstrated

  • Integration failure diagnosisTraced a silent gap between the voice platform's completion events and the pipeline to a burst of simultaneous calls, and mitigated the cause first (build timeline, 27 April).
  • Workflow designMoved recovery from a one-off script into the product: detect, fetch, replay, with an audit row for every action (architecture).
  • Idempotent processingReplay reuses the call-event dedupe key and the ordinary processing job, so a recovered event cannot create a second call record: 0 duplicate call records in 339 recovered calls.
  • Operator visibilityThe failures panel with four actions and the delivery percentage on the dashboard card, the two screenshots on this record.
  • Operational measurementThe outcome figures are read from the recovery audit table with the system's own definition of a missing event, with the limits stated beside them.

Career facts as confirmed to Colaberry; project contribution from the repository record.

What this project shows

This project shows what an AI Systems Architect is responsible for: making a failure visible, designing a repeatable response, and checking what happened after it shipped. Kes's system resolved 586 of the 604 events it detected, and the gaps it left, 18 unresolved cases and one place where the recovery path can still drift from the live one, are stated on the record rather than hidden.

Build one of these

Start the program that produced this work.

See the program