Training · learner work

The Next Call Needs a Time. And a Place in the Queue.

Campaign calls need an eligible time and a place among the other launches waiting. An AI Systems Architect at Colaberry built a shared scheduling grid and checks before launch for an admissions team's calling platform, and the operating review shows what those controls cover and where the callback path still breaks down.

verified

All student projects
When a planned call may go, and where it sits, in 135 seconds. Narrated by a synthetic voice; the figures it states are the verified metrics recorded below.
Industry
Education
Capability
Operational AI call scheduling and pacing
Status
Shipped
Built by
Colaberry team
Published
2026-09-21

The situation

Before every call an admissions team's calling system has to answer one operator's question: may this call go now, and where does it sit among the calls already waiting? A voicemail retry in two hours, a call into a campaign, a callback the lead asked for: each needs a time the lead can take it and a place among hundreds of others, because the voice platform and the people behind it can only take so many at once.

On 10 Sep 2026 one five-minute window held 14 launch jobs against a cap of four, and the commit that fixed it traced that day's stuck calls to the same root cause, which is why the choice of a slot matters as much as the choice of a time. Other CORA records cover how a call's outcome is recovered, how its intent becomes a next step and how the platform keeps work durable; this one covers only when a planned call may go and where it sits.

What it had to do

  • Give each planned call a time inside its campaign window.
  • Space launches so that no more than four are planned into any five-minute slot, as a target of the allocator rather than a cap the whole platform is held to.
  • Check pauses and guards again at the moment a call would leave.

What constrained it

  • The voice platform dials and reports calls from outside, and a burst of launches has coincided with completion reports that never arrived.
  • Leads are called by campaign: new leads any day, cold leads on weekdays only, each within its own hours.
  • Several parts of the platform plan calls: campaign entry, voicemail retries, callbacks a lead asks for, stale-lead advances and operator retries.
One schedule, six calls 75 seconds apart, and a callback sitting off the grid
One schedule, six calls 75 seconds apart, and a callback sitting off the gridEvery upcoming call by its planned time: six calls spaced 75 seconds apart, four to a five-minute slot on the grid the allocator targets, the lead requested callback off the grid at 15:02:00, a held job off the grid at 14:50:47, a cancelled retry, and its replacement on Monday. The query is a read only one typed into the console's DB Explorer by the capture script; the column names scheduled_by and on_75s_grid are the query's, not the product's. Keyless local stack at the pinned commit with invented placeholder rows, not production data; times are Central.

Decisions that made the difference

Three choices shaped the system. Each one lives at a specific point in the drawing above.

  1. At Pre-launch checks

    Check the call against a calling window

    A call at the wrong hour is a call the lead cannot take, and campaigns differ: new leads any day, cold leads on weekdays only.

    Check each launch against its campaign's configured window, 8 AM to 10 PM with cold leads on weekdays, in the zone recorded for the lead or a default when none was, and move a call that falls outside to the next opening plus an hour.

    Evidence The calling-window module and the launch job at the pinned commit, the production window settings, and a read of the zones on call events.

    4,692 of 4,743of the call events in the 30 days to 19 Sep 2026 carried a recorded Central zone; 51 carried none and fell back to the default. The launch request sends no zone, so this shows which zone the window was checked in, not where each lead is. The window is application configuration, checked only at launch and only in live mode, and a window that crosses midnight is not supported.

  2. At Take a free slot

    Make callers share one grid

    Two schedulers that read the same free slot at the same moment both take it; on 10 Sep 2026 one five-minute window held 14 launch jobs.

    Put campaign entries, voicemail retries and out-of-hours moves on one grid of 75-second slots, with four to a five-minute slot as the allocator's target, and hold a database lock while a slot is chosen.

    Evidence The allocator and its lock at the pinned commit, the burst-cap commit, and a read of launch start times by period.

    11 of 1,064five-minute clock windows since the lock in which more than four launch jobs started, by the time a worker started each call; 20 of 331 in the grid period before the lock and 60 of 2,732 on the older pacer. The periods differ in length, volume and callers, so this is an observation, not a measured effect of the lock. The lock covers only callers that go through the allocator: callbacks, stale-lead advances and the rebalancer do not, and a full four-hour search returns a slot nobody checks.

  3. At Pre-launch checks

    Recheck before the call leaves

    Between planning a call and placing it, a lead can ask for a person, opt out or enrol, and a campaign can be paused.

    Run the pauses and guards again at launch: a paused campaign holds the call, a blocked number or a do-not-call lead cancels it, and an unresolved urgent request cancels it until a sales rep records an outcome.

    Evidence The launch job and the escalation guard at the pinned commit, the guard commit, and a read of callback jobs before and after it.

    0 of 19reconstructed callback requests in the 24 Aug to 19 Sep 2026 window had a completed launch job; 18 were cancelled at or after their planned time and 1 before it. The code shows how the urgent-request guard can cancel a callback, but the platform stores no cancellation reason, so the cause of each is not recorded. In the same window all 19 requests to call later with no time, a separate category the guard does not treat as urgent, had a completed launch job. A completed launch means the voice platform accepted the request, not that the lead answered.

Who built it

  • AI Systems Architect at Colaberry: designed and built the call scheduling, the shared grid and its lock, the calling windows and the launch-time checks
A call that could not be placed yet: one row cancelled, a new row on Monday morning
A call that could not be placed yet: one row cancelled, a new row on Monday morningOne lead's history: a launch job marked completed, which means the dial request was accepted, with the call's own completion arriving separately as a Call Events row; then the retry that came due on a Saturday, cancelled, and its replacement pending at Monday 11:00:00 AM, which is 9:00 AM in the lead's own window plus the one hour buffer. Rescheduling as CORA does it: a cancel and a new row, not an edit. Keyless local stack at the pinned commit with invented placeholder rows, not production data; times are Central.

The build

  1. Callbacks at the time a lead names
  2. Calling windows by campaign
  3. Four per five-minute slot as a target
  4. One grid instead of a count
  5. A guard at launch for urgent requests
  6. A lock on choosing a slot
Notes on 6 of the 6 steps
  • Callbacks at the time a lead names. Spoken callback times are read from the call by pattern and turned into a planned time, on the UTC clock.
  • Calling windows by campaign. Campaign windows kept in the database, and a window check at launch that moves a call to the next opening.
  • Four per five-minute slot as a target. A count-based pacer spaces launches four to five minutes, after an earlier ten-per-slot setting.
  • One grid instead of a count. The count-based pacer, which assumed every launch sat on the grid it computed, replaced by a search for a genuinely free 75-second slot from a fixed starting point.
  • A guard at launch for urgent requests. A launch-time check that cancels a call to a lead whose last call carried an urgent request, a callback request among them, until a sales rep records an outcome.
  • A lock on choosing a slot. A database lock around slot allocation after a five-minute window held 14 launch jobs; the next day the guards stopped raising dashboard exceptions and log instead.
The eligibility rule the product states: days, hours and the caller's local timezone
The eligibility rule the product states: days, hours and the caller's local timezoneThe campaign settings page: New Lead every day, Cold Lead Monday to Friday, both 8 to 22, and the line "Times are in the caller's local timezone". The 2,880 minute first retry delay on the same page is why the retry above landed on a Saturday. These values are the ones the migration seeds, so they are CORA's defaults rather than settings chosen for the picture. Keyless local stack at the pinned commit with invented placeholder rows, not production data; times are Central.

What was built

A planned call is a job row with one time on it, its run time, and nothing else about when it was meant to happen. Campaign entries, voicemail retries and calls moved out of hours take that time from one shared grid: 75-second slots counted from a fixed starting point, four to a five-minute slot as the allocator's target. The allocator walks forward from the earliest allowed moment to the first slot no pending or claimed launch holds, and since 10 Sep 2026 it takes a database lock first, so two schedulers that go through it cannot both read the same free slot and fill it twice. That lock protects only the callers that go through the allocator. If four hours of slots are all taken, it returns the first slot after them without checking it, and that slot is not checked again at launch.

Stack

  • Python
  • Fastapi
  • PostgreSQL
  • Redis
  • Rq
  • Docker
More on what was built

Not every call goes through the grid. Callbacks a lead asks for, and retries after a failed booking or an unconfirmed transfer, are written at an exact time with no lock and no occupancy check, and they carry no campaign name, so a campaign pause does not hold them and a new-lead allocation can move them. Stale-lead advances still use the older count-based pacer that the grid replaced on 17 Jul 2026, and an operator retry is placed at once. A spoken clock time is read by pattern and set on the UTC clock, not the lead's, so "call me at 3pm" is planned for 10:00 AM Central Daylight Time; a callback with no time goes two hours out.

Every five minutes a rebalancer looks for any five-minute slot holding more than four pending launches. When it finds one it rewrites every pending future launch into a tight sequence starting from the earliest, four to five minutes, which keeps the cap but can pull a Monday deferral, a 48-hour retry or a callback earlier than planned; it takes no lock and leaves no record.

At launch the call is checked again. A system pause holds it; a campaign pause holds campaign calls, not callbacks; a blocked number, a lead marked do not call, a spam tag, an enrolled student or an unresolved urgent request cancels it. In live mode the call must fall inside its campaign's configured window, 8 AM to 10 PM with cold leads on weekdays, in the zone the call platform recorded on the lead's last call, or the configured default when none was recorded; outside it, the job is cancelled and a new one allocated at the next opening plus an hour, with no link between the two. The launch request itself sends no timezone, and 4,692 of 4,743 call events in the 30 days to 19 Sep 2026 carried a recorded Central zone, so the window was in practice checked in Central; that says which zone was used, not where each lead is.

A completed launch job means the voice platform accepted the request, or, in shadow mode, that the request was intercepted and no call was made. Whether the lead answered arrives later, in a separate event matched to the lead by contact and time, with no link back to the launch job.

What works: campaign calls, voicemail retries and out-of-hours moves get a place on one grid; the allocator's lock stops two schedulers that use it from taking the same slot; every call is checked again at launch against pauses, five guards and the configured window; and a completed launch job records that the voice platform accepted the request.

What remains unresolved: callbacks a lead asks for are written at an exact time with no lock and no occupancy check; a spoken time is set on the UTC clock, not the lead's; 0 of 19 reconstructed callback requests in the four weeks after the urgent-request guard shipped had a completed launch job, and the platform stores no reason for a cancellation, so the guard's part in that is an explanation the code supports and the data does not confirm; the rebalancer can pull a deferral or a callback earlier than planned; a full four-hour search returns a slot nobody checks; and no test runs two schedulers against Postgres at once.

The screenshots were captured on a local copy of the platform at the pinned commit with no keys and invented rows; the console labels Central times CST even in daylight time.

Capabilities

  • Call scheduling
  • Calling windows
  • Rate limiting
  • Pre launch guards
  • Timezone handling
  • Runtime controls

Integrations

  • Voice platform
  • Crm

Data stores

  • PostgreSQL
  • Redis
Scheduled Actions: what is coming, and when
Scheduled Actions: what is coming, and whenThe product's own scheduling page. Nine leads with campaign, last call and next action: six first touches minutes apart, a callback at ten minutes, one lead held at Now while its campaign is paused, one moved out to a day and twenty hours. The table is sorted by last call, not by next action. Keyless local stack at the pinned commit with invented placeholder rows, not production data; times are Central.

The measurement

The record measures scheduling from the platform's own tables, and says what each figure is. Since the lock, 11 of 1,064 fixed five-minute clock windows started more than four launch jobs by worker start time, against 20 of 331 before it; worker pickup delay after the final planned time was a median 16 seconds; and 0 of 19 reconstructed callback requests in the four weeks after the guard had a completed launch job. What it cannot measure is punctuality against the time a lead asked for, because that time is not kept.

Evidence maturity: a shipped scheduling layer with operating counts read on 19 Sep 2026 in read-only sessions on its own tables, and its scheduling tests run offline at the commit production runs. The hero row is empty on purpose: the pacing periods are not a controlled comparison, and the one before-and-after the record can reconstruct, callback requests around the guard, shows fewer completed launch jobs after it, with no recorded reason.

Full notes on all 6 metrics
  • 11 of 1,064 five-minute clock windows over four launch jobs since the lock, by worker start time

    Five-minute clock windows since the lock in which more than four launch jobs started, by worker start time

    verified

    Baseline
    Before the lock, on the grid: 20 of 331 windows (17 Jul to 10 Sep). Before the grid, on the count-based pacer: 60 of 2,732 (1 May to 2 Jun).
    Sample
    Every five-minute window from 10 Sep 2026, 11:14 AM, to 19 Sep 2026, 12:00 AM Central in which a worker started at least one completed launch job: 1,064 windows, 3,917 launch jobs.
    Methodology
    Read-only counts on scheduled_jobs, completed launch_outbound_call rows grouped into fixed five-minute clock windows (date_trunc to the hour plus a five-minute floor) by the time a worker claimed each job, denominator the windows with at least one such start, cut into periods at the platform's own flag history (live from 18 Apr to 10 Jul and from 15 Jul, shadow between; campaigns paused 9 Jun to 15 Jul) and at the grid commit (17 Jul 2026, 4:05 PM Central) and the lock's deploy (about 11:14 AM Central, 10 Sep 2026). Counted by planned time instead of start time, 10 windows held more than four, out of 1,064 on that clock as well, all on 10 and 11 Sep, and 57 of the 66 launch jobs in them were created after the lock. Counting only the launch jobs created after the lock, 9 of the 1,049 windows holding at least one of them held more than four of them.

    Limitations

    • Fixed clock buckets, not rolling windows: a burst that straddles a boundary can count as under four in both.
    • Worker start time, not dial time, which is not stored.
    • The periods are not a controlled comparison: volume, the parts of the platform planning calls and pauses all differ.
    • The lock covers only callers that go through the allocator; callbacks, stale-lead advances and the rebalancer do not.
    • A launch job started is a request to the voice platform, not a connected call.
  • 0 of 19 reconstructed callback requests, 24 Aug to 19 Sep 2026, had a completed launch job

    Reconstructed callback requests since the urgent-request guard with a completed launch job

    verified

    Baseline
    Before the guard, 25 Jul to 24 Aug 2026: 2 of 6 reconstructed callback requests had a completed launch job.
    Sample
    Launch jobs created from 24 Aug 2026, 8:04 PM, to 19 Sep 2026, 12:00 AM Central for a callback request (5 with no time, 14 with a time), reconstructed as one request each after 7 cancel-and-replace links were joined by identical payload within two seconds.
    Methodology
    Read-only count on scheduled_jobs, launch jobs whose payload carries a callback intent reason, with a cancelled row and a new row carrying the same payload within two seconds treated as one request moved, not two. Cancelled at or after the planned time is when the launch-time checks run. Before the guard, 25 Jul to 24 Aug 2026, 2 of 6 reconstructed callback requests had a completed launch job.

    Limitations

    • A reconstruction, not a stored record: the two-second equal-payload join is a heuristic that can merge two genuine requests or miss a move whose payload changed.
    • The platform stores no reason for a cancellation; the urgent-request guard is an explanation the code supports, not a recorded cause.
    • A completed launch job is the voice platform accepting a request, not a connected call; whether a person called the lead instead is not recorded here.
    • The 19 "call later, no time" requests in the same window are a separate category and all 19 had a completed launch job.
  • 5,247 of 5,450 launch jobs attributed to origins that normally use the shared allocator

    Launch jobs created in 30 days attributed to origins that normally use the shared allocator

    verified

    Baseline
    None: a split of one cohort.
    Sample
    Every launch job created from 20 Aug 2026 to 19 Sep 2026, 12:00 AM Central: 5,450.
    Methodology
    Read-only count on scheduled_jobs grouped by the payload keys each origin writes: campaign entry and voicemail retry go through the allocator; stale-lead advances use the older pacer; intent callbacks and retries are written at an exact time. A call moved out of hours copies its original payload, so it counts with its origin.

    Limitations

    • Origin is read from payload key names, which the code writes per origin; it is not a stored field, and it is not the execution path.
    • A call moved out of hours copies its original payload, so replacement rows count with their origin whatever path placed them.
    • Attributed to the allocator is not the same as kept to the target: see the pacing figure.
  • median 16 seconds worker pickup delay, 90th percentile 28 seconds

    Worker pickup delay after the final planned time, completed live launch jobs

    verified

    Baseline
    None: one window.
    Sample
    Every completed live launch job a worker started from 20 Aug 2026 to 19 Sep 2026, 12:00 AM Central: 4,688.
    Methodology
    Read-only percentiles over scheduled_jobs of the claim time minus the final run time for completed launch_outbound_call rows, excluding test calls.

    Limitations

    • The planned time is the FINAL one: a rebalance, a new-lead bump or a pause release can rewrite it before the call starts.
    • It is not the time a lead asked for; that time is not stored.
    • Completed live jobs only; cancelled and pending jobs are not in the sample.
  • 4,692 of 4,743 call events carried a recorded Central zone; 51 carried none and fell back

    Call events in 30 days whose recorded timezone was Central

    verified

    Baseline
    None: a split of one cohort.
    Sample
    Every call event created from 20 Aug 2026 to 19 Sep 2026, 12:00 AM Central: 4,743.
    Methodology
    Read-only count on call_events grouped by the timezone field the voice platform sends; the calling-window check reads that field from the lead's latest call event and falls back to Central.

    Limitations

    • The launch request sends no timezone, so the field is whatever the voice platform reports; whether it is where each lead is, is not established.
    • A calling window is application configuration, not a legal requirement, and nothing here claims one is met.
  • 390 of 391 scheduling tests pass offline at the pinned commit; the one failure expects a legacy zone name to resolve

    Scheduling tests passing offline at the pinned commit

    verified

    Baseline
    None.
    Sample
    The 16 test files that exercise the allocator, calling windows, pauses, guards, callbacks, voicemail retries and campaign entry: 391 tests.
    Methodology
    pytest over the 16 files in an isolated container with no network, on 19 Sep 2026. The one failure is a test that expects the legacy zone name US/Central to resolve; in the test image it falls back to Central.

    Limitations

    • Every test runs on in-memory SQLite or mocks; none runs two schedulers against Postgres at once, so the lock itself is tested only by the text of its SQL.
    • No test covers the rebalancer, the window deferral at launch, a daylight-saving change, a window that crosses midnight, or a callback reaching the launch-time guards.
    • The one failing test, test_get_contact_timezone_accepts_us_central, expects the legacy zone name US/Central to resolve; in the test image it falls back to the default. Not triaged here.

What happened next

Shipped

  • A shared grid of 75-second slots for campaign calls
  • A lock on choosing a slot
  • A calling-window check at launch
  • Pauses and guards rechecked at launch

Not pursued

  • A cap on callback retries

Unknown

  • Callbacks on the shared grid
  • A spoken time read in the lead's zone
  • Callbacks that survive the urgent-request guard
  • A rebalancer bounded to the crowded slot
  • The time a lead asked for, kept
  • A link from a launch to its completion
  • A link from a moved call to the one it replaced
Notes on 12 of the 12 items
  • A shared grid of 75-second slots for campaign calls. Campaign entries, voicemail retries and out-of-hours moves take the first free slot the allocator finds, with four to a five-minute slot as its target. When four hours of slots are full it returns the next slot unchecked, and no test runs two schedulers against Postgres at once.
  • A lock on choosing a slot. Since 10 Sep 2026, for the callers that go through the allocator only: campaign entry, voicemail retries and out-of-hours moves. Callbacks, stale-lead advances and the rebalancer take no lock.
  • A calling-window check at launch. Live mode only, in the zone the call platform recorded or a default when none was; a call outside its window is replaced by one at the next opening plus an hour, with no link to the row it replaced. No test covers the deferral, a daylight-saving change or a window that crosses midnight.
  • Pauses and guards rechecked at launch. System and campaign pauses hold calls; blocked numbers, do-not-call leads, spam tags, enrolled students and urgent requests cancel them. No test covers a callback reaching these guards.
  • A cap on callback retries. Recorded in the commit that capped the other intent retries: callback requests were left uncapped on purpose, as explicit asks from a lead.
  • Callbacks on the shared grid. Callbacks and intent retries are written at an exact time with no lock and no occupancy check, and two can share a slot; no decision about it is recorded.
  • A spoken time read in the lead's zone. Clock times are set on the UTC clock, so "3pm" is planned for 10:00 AM Central Daylight Time; the requested time is not kept anywhere, so punctuality against it cannot be measured. No decision about either is recorded.
  • Callbacks that survive the urgent-request guard. The guard treats a callback request as urgent and can cancel the callback itself at launch unless a sales outcome is recorded. 0 of 19 reconstructed callback requests in the four weeks after it shipped had a completed launch job; the platform records no cancellation reason. No decision is recorded.
  • A rebalancer bounded to the crowded slot. It rewrites every pending future launch, with no lock and no record, and can move a deferral or a callback earlier than planned; no test covers it. No decision is recorded.
  • The time a lead asked for, kept. The requested time lives only in the planned run time, which a rebalance or bump can overwrite, so punctuality against the request cannot be measured; no decision about it is recorded.
  • A link from a launch to its completion. The voice platform does not echo the launch back, so a completion is matched to the lead by contact and time; no decision about it is recorded.
  • A link from a moved call to the one it replaced. A deferral cancels one row and writes another with no link or reason, so moves cannot be counted afterwards; no decision about it is recorded.

Meet the builder

AI Systems Architect

Project contribution

Built callback scheduling in March 2026 and calling windows in April; wrote the count-based pacer, the four-per-five-minutes cap and the rebalancer at the end of April; replaced the pacer with the shared grid on 17 Jul; put campaign entry on the grid in August with a new-lead priority; added the launch-time guards from May to September; and added the lock on 10 Sep. Of the nine commits this record cites, the four from March and April 2026 name no AI coding assistant as co-author; the five from July 2026 onward name Claude.

Skills demonstrated

  • Scheduling under a shared limitA grid of 75-second slots, four to a slot as its target, and a lock on choosing one for the callers that use it (architecture, build timeline).
  • Timezones and calling windowsConfigured campaign windows checked at launch in the zone recorded for the lead or a default; the record states where that zone comes from (architecture, decisions).
  • Integration with a voice platformLaunch requests, and completions that arrive separately with no link back (architecture).
  • Verification against production dataLaunch start times by fixed clock window and period, callback requests reconstructed around a guard, each with its method stated (measurement).

From the repository record.

What this project shows

This project shows an AI Systems Architect turning 'call them again' into a time and a place: one grid for campaign calls, a lock that stops two schedulers that use it from taking the same slot, and a second look at pauses and guards before a call leaves. It also shows what scheduling does not yet hold: the callbacks leads ask for skip the grid, their spoken times are read on the wrong clock, 0 of 19 reconstructed callback requests in the four weeks after the guard shipped had a completed launch job with no recorded reason, and the rebalancer can move calls earlier than planned.

Build one of these

Start the program that produced this work.

See the program