Training · learner work
Test the Action. Before It Reaches a Customer.
A convincing message from an AI model can still be the wrong thing to send. An AI Systems Architect at Colaberry built the shadow paths in an admissions team's calling system, so that the calls, texts, emails and CRM updates it intended were held at the boundary, and the texts, emails and CRM updates it recorded could be read in full before any of them reached a customer.
verified
All student projects- Industry
- Education
- Capability
- Operational AI action gating and shadow testing
- Status
- Shipped
- Built by
- Colaberry team
- Published
- 2026-09-19
The situation
The system calls an admissions team's leads, texts and emails them, and writes what happened into their CRM record, and an AI model writes the texts and emails. Each of those actions has a consequence the words on a screen do not: a call reaches a person, a text lands on a phone, a changed CRM field changes what the admissions team does next. A model that writes a convincing message can still attach it to the wrong lead, the wrong campaign or the wrong moment.
The requirement was to let the new system run its whole decision path against real leads while holding back the actions themselves. Every job would still run and every message would still be written, but the call would not be placed, the message would not be sent and the CRM field would not change until someone opened a gate on purpose. What was held back had to be readable afterwards: the words, the fields, the lead and the campaign.
For a builder the design question was where to draw the line. Blocking the jobs would have hidden what they decide; letting them run and blocking only the last step keeps the decision path real and makes its output readable. The gate had to sit exactly where an action leaves the system, once for calls and messages and once for CRM writes, because those fail in different ways.
What it had to do
- Hold the main outbound actions and CRM writes at the point where they would leave the system, without stopping the jobs that decide them.
- Keep a readable record of each held action: the message body, the CRM fields, the lead and the campaign.
- Control calls and messages separately from CRM writes, and make every change to either an audited event.
What constrained it
- The system was replacing a working Zapier flow, so the rehearsal ran beside it on the same leads.
- Holding a call means no call happens, so the voice platform sends no completion event for it.
- The model still writes each message in shadow mode; only the sending is held.

Decisions that made the difference
Three choices shaped the system. Each one lives at a specific point in the drawing above.
- At Apply action gates
Intercept the action
A new system about to take over live calls, texts, emails and CRM updates, and no way to see what it would do short of letting it do it.
Put the check where each action leaves the system, not where it is decided: one gate inside the call and message jobs, a second inside the CRM adapter. The call, text and email jobs and the two CRM jobs read the flags from the database as they run; a few other CRM writes fall back to the environment setting. Every job keeps running; only the outward step is held.
Evidence The call gate and the CRM gate at the head commit; System Controls and the audit rows of 18 Apr 2026.
4,267intended actions for 722 leads were recorded instead of performed during the five-day rehearsal in April. The two gates then opened within seconds of each other at 12:41 AM Central on 18 Apr 2026, each change an audited row.
- At Inspect shadow records
Inspect the intended content
A record that an action was intercepted says nothing about whether the action was right. The part worth reading is what the model wrote and what the CRM would have been told.
Write the message before the gate and keep it: the whole body of every text and email, and the field map of the CRM update that follows a voicemail message, stored as shadow records and as messages with the status shadow, clearly apart from the live ones. The other CRM writes are held at the adapter and logged by field name, so they are stopped without being kept.
Evidence The commit of 5 Apr 2026 that keeps the words; the payload counts read from the shadow table.
3,204 of 4,267intended actions kept the exact words or fields they would have sent. The 1,063 calls kept their target and campaign but not their words, because the voice agent speaks from its own script during the call.
- At Verify the evidence
Separate rehearsal from proof
A rehearsal can look like proof. A held call is never placed, so nothing after it happens, and a test that is skipped can be counted as one that passed.
Say which is which. The record reports the intercepted call's missing completion event, counts a skipped evaluation as a skip, and leaves the headline row empty because nothing compares a before with an after.
Evidence The architecture note on shadow mode and the voicemail sequence; the evaluation harness run offline at the head commit, with and without its flag.
0 of 6evaluations of the call summaries and consent detection ran with a fixture, and no test judges the words the model wrote. The gates themselves pass 12 of 12 tests, against mocked services.
Who built it
- AI Systems Architect at Colaberry: designed and built the action gates, the shadow records and System Controls

The build
- Shadow logging built
- The words kept, not only the intent
- First action recorded instead of performed
- System Controls
- Both gates opened within seconds
- Shadow mode on again during a campaign pause
- Gate opened again, campaigns resumed
- A new action with its own switch
- The July switch confirmed as intentional
- Measured, read-only
Notes on 10 of the 10 steps
- Shadow logging built. One function becomes the single place a held action is written; the call, text and email jobs check the gate; 14 unit tests of the gates arrive with it, of which 12 remain at the head commit.
- The words kept, not only the intent. Texts and emails are written before the gate and stored with the status shadow; a CRM update records the fields it would have written.
- First action recorded instead of performed. The rehearsal begins on real leads at 5:28 PM Central, beside the workflow being replaced.
- System Controls. Both gates move into the database, change without a restart, and every change writes an audit row.
- Both gates opened within seconds. Calls and messages at 12:41:03.8 AM Central, CRM writes at 12:41:09.3 AM, about five and a half seconds later; the status bar reads the mode from the database six minutes later.
- Shadow mode on again during a campaign pause. Four calls recorded instead of placed in five days.
- Gate opened again, campaigns resumed. Shadow mode off at 4:44 PM Central; outbound campaigns resumed six seconds later.
- A new action with its own switch. A staff-only transcript note on the CRM record, gated by its own flag rather than either gate.
- The July switch confirmed as intentional. Reconstructed from the configuration and audit tables and recorded in the repository.
- Measured, read-only. Figures read from the system's own tables at 10:06 AM Central; the source tests run offline at the head commit.

What was built
Two gates stand at the two places where the system reaches outside itself. The first sits inside the jobs that call, text and email: with shadow mode on, the call job records the call it would have placed and returns, and the text and email jobs store the message they wrote instead of queueing it. The second sits inside the CRM adapter: with CRM writes in shadow, a field update, a task or a note returns a record of what it would have written and makes no API call. Each gate is a flag set from one System Controls page, which writes an audit row for every change. The call, text and email jobs and the two CRM jobs read the flags from the database as they run; a few other CRM calls fall back to the value in the environment file, so the page does not govern every write.
Stack
- Python
- Fastapi
- PostgreSQL
- Next.js
- TypeScript
More on what was built
What is held back is kept where the job writes it down. One function is the single write point for the shadow table, and four paths use it to record a held action: the call job, the text job, the email job and the CRM update that follows a voicemail message. Those rows carry the payload: the target and campaign of a call, the whole body of a text or an email, the field map of that CRM update. The texts and emails are also stored as messages with the status shadow, beside the live ones, which are stored as pending until the CRM delivers them. The rest is held but not kept: every other CRM write the same gate holds, a task, a note, the student summary and the contact update after a call, is returned to the job that asked for it and logged by field name, not stored.
The gates hold the action, not the thinking. The model still writes each message and the system still reads the CRM before a gate is checked, so shadow mode is not network isolation and makes the same model calls as live mode. That is deliberate: the words that would have gone out are the words a person can read. Other outward effects sit outside both gates, each on a setting of its own read from the environment rather than from System Controls: a staff-only transcript note added in July 2026 and a conversation log, both written to the CRM, and the alert emails sent to staff. A setting also names a spreadsheet mirror, but the mirror is not built and nothing reads that setting.
An intercepted call is the gap in the rehearsal. No call is placed, so the voice platform sends no completion event, and the voicemail sequence that would follow a real call does not advance; a lead stays where it was until the gate opens and a real call completes. The rehearsal tests each action, not the journey.
The dashboard reads the same tables. System Controls shows both gates and the change log; a lead's drill-down lists its shadow actions and stored messages; the pipeline trace marks each step of a lead's run as shadow and shows the payload it would have sent. How the three screenshots were made: the repository's own dashboard (dashboard-ui) was run on a laptop at commit cf233ae with no keys, so it could reach neither the voice platform nor the CRM, against placeholder rows written for the purpose: one lead, demo-01, named Demo Contact 01, whose number sits in a range that cannot be dialled. No real lead, number, message or transcript appears. The pre-flight list on System Controls reads 10 errors for the same reason, and the pipeline trace runs past the right edge of its own box, which is how that page draws it.
Capabilities
- Shadow mode
- Action gating
- Audit trail
- Operator controls
- Pre release testing
Integrations
- Voice platform
- Crm
- AI model
Data stores
- PostgreSQL
The measurement
The system recorded 4,267 intended actions during the rehearsal, 3,204 of them with the words or fields available for inspection. That shows the actions could be inspected. It does not establish that anyone reviewed those records or that the generated content was correct.
Evidence maturity: a shipped mechanism with operating counts, not a measured outcome, read on 18 Sep 2026 from the system's own tables in a read-only session and from the source tests run offline. The rehearsal recorded 4,267 intended actions for 722 leads in five days, 3,204 of them word for word or field for field. Two figures a reader might expect are not here: a before and after of the same measure, because shadow records began with this system, and a count of tested modules, which would need a change to the repository. The headline row is empty for that reason, and the gap in the evaluation tests, 0 of 6 summary evaluations run with a fixture and no test for the words the model wrote, is on the record rather than hidden.
Full notes on all 6 metrics
4,267 intended actions recorded instead of performed in five days
Intended actions recorded instead of performed during the April rehearsal
verified
- Baseline
- None. Before the commit of 28 March 2026 there was no shadow table: a job either performed its action or did not run. This sizes the rehearsal; it does not compare it with anything.
- Sample
- All shadow rows written during the April rehearsal, from the first recorded shadow action at 5:28 PM CDT on 12 April 2026 to the audited switch at 12:41 AM CDT on 18 April 2026: 4,267 rows for 722 leads.
- Methodology
- One read-only query grouping shadow_actions created before 2026-04-18 05:41:03 UTC by action_type, counting rows and distinct contacts, with the first and last timestamps. Result: CRM contact updates 1,573 for 676 leads, outbound calls 1,063 for 721, text messages 1,046 for 710, emails 585 for 585; 4,267 for 722 leads in all, first 5:28 PM CDT on 12 April and last 8:35 PM CDT on 17 April.
Limitations
- A row is an intention. The same action in live mode is also only a request: the message table has no sent or delivered status, because the CRM delivers the message.
- The count is of what the shadow table keeps, not of everything the gates hold: CRM tasks, notes, student summaries and the contact update after a call are stopped at the adapter and logged by field name, so they leave no row. Other writes and integrations sit outside both gates with checks of their own: the staff note and the conversation log in the CRM, and the alert emails sent to staff. A spreadsheet mirror is named in the settings but is not built.
3,204 of 4,267 intended actions kept the exact words or fields they would have sent
Intended actions that kept the exact words or fields they would have sent
verified
- Baseline
- None. Before the commit of 5 April 2026 a shadow row recorded that an action was intercepted, not what it would have said.
- Unit
- share of intended actions
- Sample
- 4,267 shadow rows written during the April rehearsal, from the first recorded shadow action at 5:28 PM CDT on 12 April 2026 to the audited switch at 12:41 AM CDT on 18 April 2026.
- Methodology
- One read-only query over shadow_actions in the window, by action type: rows, rows with a non-empty message_body, rows whose fields value is a non-empty object. Texts 1,046 of 1,046 with a body, emails 585 of 585, CRM updates 1,573 of 1,573 with a field map, calls 0 with either. A second query counts outbound_messages with status shadow: 1,046 texts and 585 emails.
Limitations
- A call has no content to store before it happens: the voice agent speaks from its own script. So the share can never reach every row, and the calls are the part of the rehearsal a reader cannot inspect word by word.
- Stored content is what the model wrote at the time. Nothing records whether anyone read it.
717 of the 721 leads the new system would have called were being called by the old workflow
Leads the new system would have called that the old workflow was calling in the same days
verified
- Baseline
- None. This describes how the rehearsal ran, beside the workflow it was replacing.
- Unit
- share of leads
- Sample
- 721 leads with at least one intercepted outbound call during the April rehearsal, from the first recorded shadow action at 5:28 PM CDT on 12 April 2026 to the audited switch at 12:41 AM CDT on 18 April 2026.
- Methodology
- One read-only query counting outbound call_events in the window (1,516), their distinct contacts (907), the distinct contacts with a shadowed outbound_call (721) and the intersection (717); a second counts outbound call records per Central day: 6 on Sunday 12 April, 363 on Friday 17 April, 5 on Saturday 18 April, the day of the switch, 13 on Sunday 19 April and 231 on Monday 20 April.
Limitations
- The attribution to the old workflow rests on the runbook's cutover steps and the timing, not on a field in the call record that names the caller.
- A lead on both lists was not necessarily called at the moment the new system would have called it.
13 mode changes, each recorded with its old value, new value and source
Changes to the operating mode, each recorded with what it was and what it became
verified
- Baseline
- None. Before the commit of 16 April 2026 the switches lived in the environment file, and a change left no row.
- Sample
- 13 audit rows with action mode_flag_updated, 18 April to 24 August 2026.
- Methodology
- One read-only query over audit_log where action is mode_flag_updated, ordered by time: a pause and resume at 12:23 AM CDT on 18 April; the call and message gate off and the CRM gate to live (with its shadow log switch off) within seconds of each other at 12:41 AM the same night; a system pause on 2 June and its resume on 9 June; outbound campaigns paused on 9 June; the call and message gate back on at 10:07 AM CDT on 10 July and off at 4:44 PM CDT on 15 July, campaigns resumed six seconds later; the cold lead campaign paused on 17 July and resumed on 24 August. Every row carries source system_controls.
Limitations
- An audit row records that a switch moved, not that anyone reviewed the shadow records before moving it.
- Changes made directly in the database would not pass through this endpoint; none are visible, which is not proof that none happened.
12 of 12 shadow-mode tests pass, against mocked services
Unit tests of the gates, run offline against mocked services
verified
- Baseline
- None. The tests arrived with the gates; there is no earlier version to compare.
- Unit
- share of tests
- Sample
- 12 test functions in one file.
- Methodology
- docker run --network none on a python:3.11-slim image with the project installed from pyproject.toml; pytest tests/unit/test_shadow_mode.py -v: 12 collected, 12 passed, 0 failed, 0 skipped. The Dockerfile, its ignore file and the installed package list are kept with the run.
Limitations
- A mocked test is not a customer outcome: every service in it is a stand-in.
- It says the gates behave as written at this commit, not how often they held against real traffic.
0 of 6 summary and consent evaluations ran with a fixture
Evaluations of the AI's call summaries and consent detection that ran with a fixture
verified
- Baseline
- None. The harness has had no fixtures since it was written.
- Sample
- 11 test functions: 6 marked as needing fixtures, 5 regression tests.
- Methodology
- pytest tests/evals/test_shadow_eval_harness.py -v -rs: 5 passed, 6 skipped ("Set EVAL_FIXTURES=1 and load fixture transcripts to run evals"). With EVAL_FIXTURES=1: 6 passed, 5 skipped ("fixture not loaded"); the extra pass loads no fixture. The five passing regression tests check the consent gate and summary formatting, not generated content.
Limitations
- This record does not add the fixtures. Doing so is a change to the builder's repository and a separate piece of work.
- The gap is wider than this card: no test judges the words the model wrote for a lead. The message-generator tests supply their own strings and check length, fallbacks and shape, so the texts and emails the rehearsal stored are unjudged.
What happened next
Shipped
- A gate on calls, texts and emails
- A separate gate on CRM writes
- Shadow records with the words and fields
- The message written before the gate, so the words exist to read
- System Controls with an audited change log
Paused
- Fixtures for the summary and consent evaluations
Not pursued
- A check on the words the model writes to a lead
- A rehearsal of the whole journey
- A recorded review of the shadow records
- The staff note and conversation log behind the same gates
- A spreadsheet mirror
Notes on 7 of the 11 items
- The message written before the gate, so the words exist to read. The text and email jobs call the model first and check the gate afterwards, which is why a held message still has its exact words; it is also why shadow mode makes the same model requests as live mode.
- Fixtures for the summary and consent evaluations. Started, not finished: the harness exists and its fixture list is empty, so five of its six evaluations skip without running. Adding them is a change to the builder's repository and was not made for this record.
- A check on the words the model writes to a lead. Proposed, not built: the rehearsal stores every word the model wrote, and no test judges those words. The message-generator tests supply their own strings and check length, fallbacks and shape; the evaluation harness covers call summaries and consent detection.
- A rehearsal of the whole journey. Not built: the current shadow implementation does not simulate the completion callback, so it does not rehearse the full journey. An injected completion event could exercise the steps after a call, but it would not show that a real call was delivered.
- A recorded review of the shadow records. Proposed, not built: nothing records that a person read a shadow record, judged it or changed something before the gates opened.
- The staff note and conversation log behind the same gates. Kept separate by design, each on a setting of its own read from the environment rather than from System Controls; the two gates therefore do not cover every outward effect.
- A spreadsheet mirror. Named in the settings and not built: the adapter is a stub, and nothing reads the setting that would govern it.
Meet the builder
AI Systems Architect
Project contribution
Built the shadow logging and the gate on calls, texts and emails on 28 Mar 2026; kept the generated words and CRM field maps in the shadow records on 5 Apr; moved both gates into the database behind an audited System Controls page on 16 Apr; both gates were opened on 18 Apr, which the audit rows record without naming who moved them. In July added a staff-only transcript note behind a switch of its own, and confirmed the July use of shadow mode from the system's own tables. The three commits that built the gates carry no co-author; the ones from 18 Apr onward name an AI coding assistant as co-author, which the record states rather than hides.
Skills demonstrated
- Choosing where to interceptPut each gate at the point the system reaches outside itself, inside the call and message jobs and inside the CRM adapter, so every job still runs (architecture).
- Keeping what was held back readable3,204 of 4,267 intended actions kept the exact words or fields they would have sent (measurement).
- Separate controls for separate failuresA call and a CRM write fail in different ways, so each has its own gate, and both open from one audited page (architecture, build timeline).
- Naming what a rehearsal cannot showNames the gaps rather than hiding them: the missing completion event, the empty evaluation fixtures and the words no test judges (roadmap, measurement).
From the repository record.
What this project shows
This project shows an AI Systems Architect deciding where an AI system may act and where it may only propose: two gates at the boundary, selected actions kept for reading, and every change of mode on the record. It ran for five days on real leads before anything reached them from the new system, and 3,204 of the 4,267 actions it held kept their words or fields. It also shows the discipline of saying what that does not prove: no review is on record, the rehearsal stops at a held call, and no test judges the words the model wrote.
Build one of these
