Training · learner work

The Workflow Grew. So Did the Responsibility.

A recap workflow begins with a call and ends with a useful next step. As an admissions team's calling system took on more of that work, the harder question became where unfinished work lives and who moves it forward. An AI Systems Architect at Colaberry built the answer into a Python API and worker platform, with explicit boundaries between process state, execution and the CRM, and the record states where those boundaries still leak.

verified

All student projects
Who owns a job, from the call to the CRM, in 115 seconds. Narrated by a synthetic voice; the figures it states are the verified metrics recorded below.
Industry
Education
Capability
Operational AI platform state and execution
Status
Shipped
Built by
Colaberry team
Published
2026-09-19

The situation

When a call ends, an admissions team's calling system has work to finish: analyse the call, move the lead, schedule the next contact and update the CRM. Each step can fail, and a step that is lost leaves a lead waiting with nobody told.

The platform's design documents state what it was built to replace: three Zapier workflows, for inbound recap, outbound cold-lead recap and outbound new-lead recap. The first architecture decision, on 8 Mar 2026, chose one codebase for an API and its workers and named what that had to make straightforward: a dashboard and replaying work. Nothing in the repository records those workflows being retired, so this record describes what the platform took responsibility for, not a finished migration.

What it had to do

  • Keep every piece of unfinished work somewhere that survives a restart.
  • Answer the systems that call in at once, and do the work separately.
  • Send writes to the CRM through clients that can be switched off, retried within a bound, and seen when they fail.

What constrained it

  • The CRM owns contacts, tasks and notes, and acts on its own: it re-enrols contacts the platform has flagged, so it is not downstream of the platform.
  • Redis runs as a broker with no persistence, so a queue alone cannot be the record of what is left to do.
  • The voice platform dials and reports calls from outside; a completion that never arrives has to be found, not waited for.
Exceptions Monitor: open exceptions sorted by the boundary they crossed
Exceptions Monitor: open exceptions sorted by the boundary they crossedThe console's open exceptions, each tagged CRM, system, call or AI, with a critical or warning level, and the actions an operator can take on each: resolve, ignore or investigate. A pending-call exception is also closed by the system when its call ends. Captured on a local copy at the pinned commit with no keys, invented rows and the default shadow mode; production runs live.

Decisions that made the difference

Three choices shaped the system. Each one lives at a specific point in the drawing above.

  1. At Write the job row

    Give unfinished work a durable home

    A job that exists only in a queue is gone when the queue is, and this Redis keeps no copy.

    Write every job as a row in Postgres first and queue it second, and let a loop every 30 seconds queue any row that is due and not running.

    Evidence The scheduler, the session helper and the recovery loop at the pinned commit, and a read of the jobs table.

    0 of 349open jobs were overdue or holding an expired lease at 8:00 AM on 19 Sep 2026. The opposite window is open: a job queued whose row then fails to commit is dropped when a worker claims it, and nothing looks for it.

  2. At Claim with a lease

    Separate accepting work from doing it

    A caller needs an answer at once, while the work it starts can take minutes and can fail halfway.

    Answer as soon as the job row is committed, let a worker claim it with a five-minute lease, and return any expired lease to pending on the next pass of the loop.

    Evidence The webhook route, the claim and lease code and the four durability test files at the pinned commit.

    82 of 82durability tests pass offline. Recovery still runs on one worker role, and an API that starts while Redis is down records work but queues none of it until it restarts.

  3. At Write to the CRM

    Make external writes an explicit responsibility

    The CRM owns contacts, tasks and notes and acts on its own, so a write the platform repeats can reach a person twice.

    Send CRM writes through clients that each carry a bounded retry and a write switch, check the operating modes on every job, and open an exception when a job fails for good, for an operator to retry, cancel or finish.

    Evidence The three CRM clients, the mode flags and the exception routes at the pinned commit, and a read of the exceptions table.

    0 of 40exceptions opened in the 30 days to 19 Sep 2026 are still open: 27 resolved, 24 of them by the system when a call finished, and 13 ignored. A task write carries no idempotency key, so the check before it is the only guard, and a contact update simply overwrites.

Who built it

  • AI Systems Architect at Colaberry: designed and built the job model, the queue and workers, the recovery loop and the CRM clients
Queue Health: work the loop recovers, and work an operator cancels
Queue Health: work the loop recovers, and work an operator cancelsCropped to two tables: jobs stuck past their time, each with a Cancel action for an operator, and expired leases, which the worker returns to pending by itself while the default worker is running. Captured on a local copy at the pinned commit with no keys, invented rows and the default shadow mode; production runs live.

The build

  1. Tests before the work was wired
  2. The webhook schedules durable work
  3. One deployment for API, Redis and workers
  4. The loop that re-queues due work
  5. Operating modes in the database
  6. An operator can advance a stuck lead
  7. A sweeper for missing completions
  8. Call launches through the platform
  9. Every uncapped retry loop capped
  10. Stuck calls recovered
  11. A job that retried forever
  12. CRM timeouts retried
  13. The commit production runs
Notes on 13 of the 13 steps
  • Tests before the work was wired. Unit tests for claims, exceptions, retries and scheduling arrive a day before the call webhook starts scheduling jobs.
  • The webhook schedules durable work. The call webhook writes a job; the API binds a Redis queue at startup.
  • One deployment for API, Redis and workers. A compose file, a Dockerfile and the service topology.
  • The loop that re-queues due work. A scan of Postgres for due pending jobs, on an interval.
  • Operating modes in the database. Flags read per job from a table that overrides the environment, audited, drawn in System Controls.
  • An operator can advance a stuck lead. A manual recovery action for work that has stopped moving.
  • A sweeper for missing completions. Calls whose completion never arrived are found and recovered or rescheduled.
  • Call launches through the platform. The CRM launches calls through the platform's own intake.
  • Every uncapped retry loop capped. Loops that could retry without a bound were capped or removed.
  • Stuck calls recovered. The sweeper extends to calls stuck mid-flight.
  • A job that retried forever. A per-job timeout, after a long job was killed mid-run and retried from scratch each time its lease expired.
  • CRM timeouts retried. A timeout the CRM reports as an authorisation failure is retried, and doomed writes are guarded.
  • The commit production runs. A failed session is rolled back before its job is failed or rescheduled.

What was built

Three systems own different parts of the work. Postgres holds process state: a row for every scheduled job, the lead, each call event and each exception, across 23 migrations. Redis and its queue carry execution only; the deployment file runs Redis as an ephemeral broker because job state lives in Postgres. The CRM owns contacts, tasks and notes, and the platform writes to it through three clients: one for contact records (their fields, tasks, notes and tags), one for call notes and one for the call log. A fourth owner is runtime configuration: operating modes live in a database table that overrides the environment and is read on every job, so a change takes effect without a restart.

Stack

  • Python
  • Fastapi
  • PostgreSQL
  • Redis
  • Rq
  • Docker
More on what was built

A job is accepted before it is done. The call webhook refuses a payload with no call id; otherwise it writes the job row and enqueues it inside one database session, and answers 202 Accepted once that session has committed. If the queue refuses, the row commits as pending and a loop that runs every 30 seconds enqueues it. A worker claims a job with a five-minute lease; each pass of the loop first returns expired leases to pending, then queues whatever is due. A five-minute sweeper looks for calls whose completion never arrived.

Three limits sit on that path. The job is enqueued before its row commits, so if the commit fails the queue holds a job with no row; the worker that claims it logs "job not found" and drops it, nothing sweeps for it, and no test exercises it. Only the default worker role runs the recovery loop, a choice the deployment file records to avoid duplicate enqueues, so while that worker is down nothing is recovered. And the API binds its queue once at startup: if Redis is down at that moment, every webhook writes its row and queues nothing until the API restarts.

Writes to the CRM are bounded and gated. Each of the three clients retries a failed call up to a configured maximum of three with a doubling delay. The contact-record client refuses a write unless CRM writes are enabled, read from the operating modes when the job passes them and from settings otherwise; the note and call-log clients have switches of their own, environment settings that are off by default in code. A task write carries no idempotency key: the client leaves the duplicate check to its callers. A contact update is a plain overwrite, so the last write wins. Call events are protected by a unique key per call and action. When a job fails for good it usually opens an exception, and the alerting service emails an alert when it first sees a new one open; two job types, the nurture scheduler and the slot rebalancer, fail without opening one, though their failures still count toward an error-rate alert. A pending-call exception is closed by the system once its call reaches a final status; for the rest, an operator can retry the job now or later, cancel its future jobs or force it to finish.

In code the default posture is shadow: outbound actions are intercepted and CRM writes are logged, not sent, and the README presents shadow as the default. Production does not run that way: its database flags have kept shadow mode off since 15 Jul 2026 and CRM writes live since 18 Apr 2026. Two parts of the design documentation do not match the running system either: a retry policy module is declared and nothing uses it, and a worker runs for a retries queue whose job function does not exist.

The two screenshots were captured on a local copy of the pinned commit with no keys, invented rows and the default shadow mode, not on production data.

Capabilities

  • Durable job state
  • Queue execution
  • Bounded retries
  • Exception handling
  • Crm integration
  • Runtime controls

Integrations

  • Voice platform
  • Crm
  • Email alerts

Data stores

  • PostgreSQL
  • Redis

The measurement

The record shows where responsibility for unfinished work sits and how it fails, not a before and after: the earlier workflows left nothing in this repository to measure against. In the 30 days to 19 Sep 2026 the platform accepted 83,604 jobs and 82,701 completed; at 8:00 AM that day none of its 349 open jobs was overdue or holding an expired lease, and none of the 40 exceptions opened in the month was still open, though the system had closed 24 of them itself and 13 were ignored. None of that measures whether a lead was served well, and a job the queue holds without a row cannot be counted at all.

Evidence maturity: a shipped platform with operating counts, read on 19 Sep 2026 in a read-only session on its own tables, and its durability tests run offline at the commit production runs. No figure compares a before with an after, and none measures an outcome for a lead.

Full notes on all 4 metrics
  • 82,701 of 83,604 jobs completed

    Jobs accepted in the 30 days to 19 Sep 2026 that ran to completion

    verified

    Baseline
    None. The earlier workflows left no record in this repository that could be counted the same way.
    Sample
    Every job created from 20 Aug 2026, 8:00 AM, to 19 Sep 2026, 8:00 AM Central: 83,604.
    Methodology
    Read-only count on the scheduled_jobs table grouped by status, for rows whose created_at falls in the 30 days before the read. The largest types in the window are metrics collection, the nurture scheduler, call-slot rebalancing and the webhook recovery sweeper, each about 8,000 to 28,000 runs, then outbound call launches (5,450, of which 434 were cancelled and 339 were still open) and call-event processing (5,299).

    Limitations

    • Completed is the job function returning, not the outcome the job was for.
    • More than half the jobs are the platform's own periodic work, such as metrics collection and the recovery sweeper, so the count says how much the platform ran, not how many leads it served.
    • A queued job whose row never committed leaves no row, so it cannot appear in this count.
  • 0 of 349 open jobs overdue or holding an expired lease

    Open jobs overdue or holding an expired lease, at 8:00 AM on 19 Sep 2026

    verified

    Baseline
    None: a single reading.
    Sample
    The 349 open jobs at 8:00 AM Central on 19 Sep 2026.
    Methodology
    Two read-only counts on scheduled_jobs. The loop runs every 30 seconds and resets expired leases first, so a working loop leaves no pending job overdue by five minutes and no expired lease.

    Limitations

    • One reading, taken while the default worker was running.
    • Nothing records a recovery: the loop resets the job row and logs a warning, so recovered work cannot be counted afterwards.
  • 0 of 40 exceptions still open

    Exceptions opened in the 30 days to 19 Sep 2026 that are still open

    verified

    Baseline
    None: a single reading.
    Sample
    40 exceptions: 25 for a call still pending, 8 for an outbound call held back by an urgent escalation, 3 for a blocked number, 2 for a lead marked do not call, 1 failed CRM message update and 1 failed staff note.
    Methodology
    A read-only count on the exceptions table grouped by type and status. Resolved and ignored are the only statuses besides open; across all time the table holds 2,914 exceptions, 37 resolved and 2,877 ignored, none open. A second read-only count at 8:14 AM, over the same window, split the 27 resolutions by whether the recorded actor is the system; no other actor was selected.

    Limitations

    • Ignored is a disposition, not a fix, and it is the larger share across all time.
    • Resolved does not mean a person acted: 24 of the 27 were closed by the system when a pending call finished.
    • A failure that never raised an exception is not counted: a queued job with no row is dropped with a log line only, and the nurture scheduler and the slot rebalancer fail a job without opening one, though those failures still count toward the error-rate alert.
  • 82 of 82 durability tests pass, offline

    Tests passing in the four durability test files

    verified

    Baseline
    None.
    Sample
    82 tests.
    Methodology
    pytest over test_worker_scheduler (18), test_worker_claim (20), test_worker_retry (21) and test_webhook_recovery_jobs (23): 82 passed, 0 failed, 0 skipped.

    Limitations

    • A test against a mocked database shows the logic, not what happens on real data or under a real outage.

What happened next

Shipped

  • Accepting a call event before doing its work
  • A job row for every piece of unfinished work
  • Job state kept out of the queue
  • Leases that return abandoned work to pending
  • A loop that queues due work every 30 seconds
  • A sweeper for calls whose completion never arrived
  • Operating modes read from the database on every job
  • CRM writes through three clients
  • Pending-call exceptions closed when the call ends
  • Operator actions on an exception
  • A bound on every retry loop

Not pursued

  • Recovery on more than one worker role

Unknown

  • An exception for every job that fails for good
  • A sweep for a queued job whose row never committed
  • The API re-binding its queue after Redis returns
  • The documented retry policy in use
  • A job function for the retries queue
  • An idempotency key on CRM task writes
  • Retiring the earlier Zapier workflows
Notes on 19 of the 19 items
  • Accepting a call event before doing its work. The webhook writes the job and answers 202 Accepted once the session commits; a payload with no call id is refused.
  • A job row for every piece of unfinished work. Written before the job is queued; a queue failure leaves it pending for the loop.
  • Job state kept out of the queue. Redis runs with no persistence; the loop rebuilds a lost queue from the rows in Postgres.
  • Leases that return abandoned work to pending. Five minutes, reset on each pass of the loop.
  • A loop that queues due work every 30 seconds. Each pass returns expired leases to pending first, then queues every due row.
  • A sweeper for calls whose completion never arrived. Every five minutes, including calls stuck mid-flight.
  • Operating modes read from the database on every job. A change takes effect on the next job, without a restart.
  • CRM writes through three clients. Contact records (fields, tasks, notes and tags) through one; call notes and the call log through two more, each with its own credentials and write switch, all under the same retry bound.
  • Pending-call exceptions closed when the call ends. The system resolves them itself once the call reaches a final status; no person acts.
  • Operator actions on an exception. Retry now, retry later, cancel future jobs, force a finish.
  • A bound on every retry loop. The last uncapped loops were capped on 24 Aug 2026.
  • Recovery on more than one worker role. The deployment file records the choice: only the default worker runs the loop, so nothing is queued twice.
  • An exception for every job that fails for good. Sixteen of the eighteen places that fail a job open one; the nurture scheduler and the slot rebalancer do not, though their failures count toward the error-rate alert, and no decision about them is recorded.
  • A sweep for a queued job whose row never committed. Nothing sweeps for it and no test exercises it; no decision about it is recorded.
  • The API re-binding its queue after Redis returns. Today a restart is needed; the fallback is described in the startup code, and no change to it is recorded.
  • The documented retry policy in use. Declared and tested, and imported by nothing in the running system.
  • A job function for the retries queue. A worker runs for the queue; the function it names does not exist.
  • An idempotency key on CRM task writes. The adapter leaves the duplicate check to its callers.
  • Retiring the earlier Zapier workflows. The design documents state the intent to replace them; nothing in the repository records a cutover, so whether any were switched off is not known.

Meet the builder

AI Systems Architect

Project contribution

Built the job row, the queue and the deployment in March 2026; moved operating modes into the database in April; added manual and automatic recovery from April to September; routed call launches through the platform in August; capped every uncapped retry loop; and fixed a timeout and a failed-session bug in September. Commits from May 2026 on name an AI coding assistant as co-author; the earlier ones name none.

Skills demonstrated

  • Deciding where state livesEvery job is a row in Postgres before it is queued; Redis is kept as a broker with no copy (architecture).
  • Designing the failure pathLeases, a 30-second recovery loop and a five-minute sweeper (architecture, build timeline).
  • Bounding retriesThe last uncapped retry loops capped on 24 Aug 2026; CRM calls bounded by a configured maximum (build timeline, architecture).
  • Naming the gapsThe record states the dropped job, the single recovery role and the unused retry policy (roadmap).

From the repository record.

What this project shows

This project shows an AI Systems Architect deciding where unfinished work lives and who moves it forward: a row in Postgres before anything is queued, a lease that returns abandoned work to pending, and a bounded retry and a write switch on every client that writes to the CRM. In the 30 days to 19 Sep 2026 the platform accepted 83,604 jobs and 82,701 completed. It also shows where that responsibility still leaks: a queued job whose row never commits is dropped, recovery runs on one worker role, and a CRM task write carries no idempotency key.

Build one of these

Start the program that produced this work.

See the program