Training · learner work
The Request Was Messy. The Price Couldn't Be.
Reservation requests arrived as web forms, forwarded threads and free text. Colaberry's Managing Director, working as its AI Systems Architect, built a queue that extracts trip details, calculates prices through explicit rules and prepares a reply for a person to review. The calculation is repeatable for the same inputs and rate-card version; the final reply still needs checking, because its model-written amount is not validated against that calculation on the queue path. Sunset in September 2026; told here as a retrospective, limits included.
verified
All student projects- Organisation
- an executive ground-transportation company
- Industry
- Transportation
- Capability
- Operational AI reservation intake deterministic pricing and a held or ready quote queue
- Status
- Paused
- Built by
- Colaberry team
- Published
- 2026-09-22
The situation
At an executive ground-transportation company, a reservation request could arrive three ways: the vendor's web form sent a structured notification, a colleague forwarded a thread, or a customer wrote a few lines. The person answering had to find the trip inside the message, price it from a rate card that lived in a document, and reply. The question the platform had to answer was not only what the price is but whether the price can be trusted, and when a quote may be called ready without a person reading it first.
On 22 Jun 2026 the COO sent four real inbound examples with the instruction to respond to any of them: the vendor's notification, a second web-form vendor's notification, a sales-lead form and a direct free-form email. The same day's notes record that the live booking mailbox was mostly free-form threads rather than structured requests. That made the reading the hard part and the arithmetic the part that could not be allowed to move.
Picture one request, an invented example rather than a recorded one: a forwarded thread asking for a car from the airport to a downtown hotel next Tuesday for three passengers. The vendor-form parser does not recognise it, so a model fills only the fields it may: addresses, date, passengers, service type. The market comes from the pickup city and the route matches a flat rate; the engine prices it line by line; because a model read the fields, the status rule routes it for review. A person opens the row, reads the line items and the warnings, has a reply drafted by a model, checks its amount against the line items by eye, and sends by hand, a send that is dry unless the delivery flag is set.
What it had to do
- Read the vendor's form by rules and free text by a model that only fills fields; a general-inbox email with none of the forty transportation terms is never read, and a long thread keeps only its first 14,000 characters.
- Price the trip as a pure function of typed inputs, so the same trip prices the same way every time, and show the arithmetic.
- Label a quote ready only when no model-read field and no warning that needs a person touched it, at a rule-assigned 0.9 for a flat-rate route and 0.7 for any other clean priced quote; route everything else for a person's review in a queue that corrects itself.
What constrained it
- Requests arrived in three shapes: a vendor web form, forwarded threads, and free text; most live mail was free-form.
- The rate card was a document with seven markets, flat routes, customer categories, minimums and surcharges, corrected after a sync with the client in May.
- A wrong number sent to a customer is worse than a slow one, so nothing could leave without a person unless the platform could show why the quote was safe.

Decisions that made the difference
Three choices shaped the system. Each one lives at a specific point in the drawing above.
- At Price, a pure function
Separate the reading from the arithmetic
Requests arrived as a vendor form, a forwarded thread or prose, and a model asked to write a reply had produced generic text with no real price; a number a model invents cannot be checked.
Read the vendor form by rules and free text by a model that may only fill fields, then price the trip as a pure function of typed inputs with every rate a literal in one file; mark any quote built from a model-read field as one a person must review; hold a request the engine cannot price or route for a person rather than guess.
Evidence The pricing engine, the vendor parser, the extractor and the status rule at the pinned commit, and the commits that introduced them.
239 of 239tests pass at the last commit across the nine reservation and pricing suites, the pricing suite among them with the test that the same inputs return the same quote; the ingest loop, the merge, the send, the drafter and the extractor have no test. The price is deterministic given its inputs; for free text the inputs are a model's reading, and most of the rate card (dead legs, hours, tolls, overnight) is unreachable from a mailbox because the ingest never sets those fields.
- At Review or ready label
Explain the number before trusting it
A quote that only shows a total cannot be checked by the person about to send it, and a status called ready is a promise the platform has to be able to keep.
Show every quote as its line items with running totals, warnings and a margin band, in the queue and in a Quote Tester; label a quote ready at a rule-assigned 0.9 for a flat-rate route with no warning that needs a person and at 0.7 for any other clean priced quote, route everything else for review, and treat the values as labels rather than probabilities; write the 0.90 auto-send threshold as a constant and a helper, and ship nothing that calls it.
Evidence The status rule and the gate at the pinned commit, the Quote Tester and the queue page, and the trust commit of 20 Jun 2026.
0 callersthe identified send route requires a person's action and is dry unless the delivery flag enables it, and the helper has no caller, so the person's action is the send boundary at the last commit; how many quotes were delivered is not in the tracked record. The customer-facing reply on the queue path is still model-written and graded by a rubric that passes any dollar figure; the deterministic guard covers the older flow only.
- At Auto-merge of duplicates
Make the queue self-correcting without sending anything
The same request arrived twice from two forwarders, noise landed beside real requests, and a scan of two general inboxes re-ran a paid extraction on every non-booking email every ten minutes for a week.
Merge duplicates persistently by reservation number, trip signature or sender and subject; let one click on a misfiled row teach a rule for that sender; set the lifecycle from who wrote last; and put a 40-term pre-filter and a seen-set in front of the model on general inboxes.
Evidence The merge, the classifier rules, the lifecycle reader and the cost guard at the pinned commit, and the commits of 23 and 29 Jun 2026.
13 rows mergedrows were merged into 11 groups by the first backfill on 24 Jun 2026, one run with no later count and no proof that every row was a duplicate. The last commit put a 40-term pre-filter and a process-local seen-set in front of the extractor after 19,809 extraction calls had run against 191 rows, about 100 dollars in eight days, two unlike counts from one investigation; the seen-set clears on every restart and no post-fix figure exists in the tracked record.
Who built it
- Managing Director and AI Systems Architect at Colaberry: designed and built the reservation intake, its deterministic pricing engine, its ready-or-held status rule and its self-correcting queue

The build
- The pricing engine and the vendor-form parser
- Five corrections after a sync with the client
- The Quote Tester
- A deterministic guard over a model-written quote body
- The booking mailbox ingested around the clock; the 0.90 gate written
- Every inbound channel in one queue; two general inboxes scanned
- Duplicate forwards merged persistently
- The last commit: the extraction re-loop closed
- Held at draft-only; a rate-rule question open
Notes on 9 of the 9 steps
- The pricing engine and the vendor-form parser. A quote calculator over seven markets with the rate card as literals, and a regular-expression parser for the vendor web form's notification emails; the first 37 pricing and 13 parser tests.
- Five corrections after a sync with the client. One trip fee per booking, flat rates for airport and golf-network routes only, per-route tolls, billed versus actual mileage in the labels, and the dead-leg policy; 62 of 62 tests at that commit.
- The Quote Tester. Paste a vendor email or enter a trip by hand, run the same engine, and read how the quote was built line by line with running totals and a margin badge.
- A deterministic guard over a model-written quote body. A validator that rejects a model-written reply unless it carries the exact total, the customer's first name and both cities, and no forbidden promise; on the older inbound flow.
- The booking mailbox ingested around the clock; the 0.90 gate written. Pull, price and persist every ten minutes, a status per row, the first queue endpoints; the same day an auto-send threshold as a constant and a helper that nothing calls.
- Every inbound channel in one queue; two general inboxes scanned. The older Inbound tab retired, the Tester nested under Reservations, and an onlyBookings mode that keeps requests and drops noise from two people's inboxes; the road-miles key was configured on the deployed platform the same day.
- Duplicate forwards merged persistently. A dedup key by reservation number, trip signature or sender and subject, and a merge that runs every ingest cycle and once as a backfill: 13 rows merged into 11 groups.
- The last commit: the extraction re-loop closed. A 40-term pre-filter and a process-local seen-set in front of the model on general inboxes, after 19,809 extraction calls against 191 rows in about eight days; no later commit exists.
- Held at draft-only; a rate-rule question open. Two days after the last commit, correspondence between the client and the builder records the tool pulling, pricing and drafting quotes but held at draft-only pending the reservation desk lead's sign-off and one open question, whether a round-trip base-rate rule written for one customer category applies to the others; a plan to start auto-sending high-confidence quotes was set for the next day's meeting, whose automated summary of 2 Jul records pricing adjustments and the quoting process discussed and no decision. No commit followed; the last commit is 29 Jun 2026, so at every commit the only send is a person's click.

What was built
The workflow, as the code implements it at the last commit. Every ten minutes the platform pulls the booking mailbox and two general inboxes through Microsoft Graph with a 24-hour lookback. On a general inbox a 40-term pre-filter and a process-local set of message ids already judged not a booking run before any model call; the booking mailbox is never pre-filtered. A message in the vendor's web-form format is read by regular expressions: passenger, date, service type, reservation number, passengers, vehicle, pickup and dropoff. Anything else is read by a model at temperature zero, which may supply only addresses, service type, passengers, date, name, vehicle and notes, and returns nothing at all without a key. The two readings are merged field by field, the rigid parse first. The market comes from the office header or the pickup city, the customer category from the sender's domain, the road miles from a distance lookup when one is possible, and the payment method is credit card by construction.
Stack
- TypeScript
- Node.js
- Express
- Next.js
- PostgreSQL
- Sequelize
- Docker
More on what was built
The price is a pure function of those typed inputs. Seven markets, a per-mile rate with a 200-mile minimum, twenty-five flat routes, customer categories, an Iowa tax that applies only when every stop is in Iowa, a fuel surcharge and a card fee are literals in one file that reads no environment variable, no settings table and no database. A status rule then decides: forward for the forward-only market; needs review for anything a model read; ready at 0.9 for a flat-rate quote with no warning that needs a person, ready at 0.7 for any other clean priced quote; needs review otherwise, and the 200-mile minimum warning itself needs a person, so a distance trip under 200 miles is always held. A person works the queue: the parsed trip, the line items, the warnings, a reply drafted by a model in the account's own voice, and a Send that prepares rather than sends unless a flag is set. Duplicate forwards are merged persistently by reservation number, trip signature or sender and subject, and a one-click reclassification teaches the queue a rule for that sender.
What works, on the evidence: the arithmetic is repeatable for identical inputs and rate-card version, and a test asserts that the same inputs return the same quote (239 of 239 tests pass at the last commit across the nine reservation and pricing suites, in 173 seconds on 21 Sep 2026); a model-read field always marks the quote for a person; the Quote Tester shows every line of a quote with its running total, its margin band and its warnings, so a person can see why a number is what it is; duplicates from two forwarders collapse into one row; and the extraction re-loop that cost about 100 dollars in eight days was found from the bill and answered in the last commit by a pre-filter and a seen-set, with no measured after.
What remains unresolved, on the same evidence. Two words in the product mean less than they sound: an auto-ready status at 0.7 sits below the only gate, and the 0.90 gate is a constant and a helper that no production code calls, so the only send route at the last commit requires a person's action and is dry unless a delivery flag is set, and the record does not establish how many quotes were delivered; the product's own dashboard says auto-send stays off. The customer-facing reply on the queue path is written by a model and graded by a rubric that passes any dollar figure; the deterministic guard that checks the exact total, the name and the cities covers only the older inbound flow. Most of the rate card is unreachable from a mailbox: the ingest never sets dead-leg miles, hours, tolls, overnight nights or per-diem days, and vehicle class and passenger count are not price inputs. Distance was blank for the queue's first sixteen days, until the road-miles key was configured on the deployed platform on 22 Jun 2026, so a distance quote in that window carried zero actual miles, was billed at the 200-mile minimum with a warning naming the zero (row f-mileage-minimum), and was routed for a person's review by that warning; what was sent in that window is not in the record. Messy requests are dropped as well as parsed: a general-inbox email with none of the forty terms is never seen, and a question that matches a FAQ entry is never extracted. A long thread keeps its first 14,000 characters, so the newest messages are the ones lost. The send has no idempotency key. The cost guard is a process-local set cleared on every restart, with no post-fix figure anywhere in the tracked record. Computed dead-leg pricing, reply-thread stripping, forward merging and round-trip pairing exist only in a local batch that was never committed or deployed. The reservation tables themselves have no migration in the repository; production created them by hand.
The screenshots were captured on a local copy of the platform at the pinned commit with no keys and invented requests the product's own code priced; the times they show are Central.
Capabilities
- Reservation intake
- Vendor form parsing
- Deterministic pricing
- Quote triage
- Human review
- Duplicate merge
Integrations
- Microsoft graph
- A language model API
- A road distance API
- A reservations web form vendor upstream
Data stores
- PostgreSQL

The measurement
The record measures what the repository can show. The episode: 19,809 extraction calls against 191 rows, about 100 dollars in eight days, 95 percent of model spend, answered on 29 Jun 2026 by a pre-filter and a seen-set, with no measured after. The queue: 43 rows in four buckets on 23 Jun; 13 rows merged into 11 groups by the first backfill. Verification: 239 of 239 tests at the last commit across nine suites. Each figure carries its date and its limit.
Evidence maturity: tracked counts and one measured cost episode; no production database was read, the browser login flow was never verified, and the send flag's production state is not in the record.
Full notes on all 4 metrics
19,809 extraction calls against 191 reservation rows in about eight days, to 29 Jun 2026
Paid extraction calls against reservation rows, the June 2026 episode
verified
- Baseline
- 191 reservation rows persisted in the same period, the count the record sets the calls against.
- Sample
- Every call the reservation ingest made to the paid trip extractor on the two general inboxes over about eight days in June 2026, as the tracked record counted them on 29 Jun.
- Methodology
- Read from the tracked record at the last commit and from that commit's own message, which agree.
Limitations
- A count of calls, not of requests: the record does not say how many of the 191 rows were quote requests.
- No post-fix call rate or spend exists anywhere in the tracked record; the run-rate is an estimate, not a saving.
- The exact eight-day window and whether the 24-hour or 72-hour lookback ran are not established.
10 of 43 rows in the queue on 23 Jun 2026 needed a reply; 10 were awaiting the customer, 19 were not a quote, 4 were marked completed
Rows needing a reply in the queue on 23 Jun 2026, of all rows that day
verified
- Baseline
- None: one snapshot on one day.
- Sample
- The rows the queue query saw on 23 Jun 2026 after a backfill, as the tracked record lists them by bucket.
- Methodology
- Read from the tracked record at the last commit.
Limitations
- A snapshot, not a rate or a trend.
- The buckets were redefined between the snapshots of 22 and 23 June, so the snapshots are not comparable with each other.
- Nineteen of the 43 rows were noise the booking mailbox persisted and then classified out.
- Completed is a lifecycle label set from who wrote last; it is not four bookings and not four delivered quotes.
13 rows merged into 11 groups by the first auto-merge backfill, 24 Jun 2026
Rows the first merge backfill combined, 24 Jun 2026
verified
- Baseline
- None: one run.
- Sample
- The rows the persistent auto-merge absorbed on its first run over the existing queue, as the tracked record states.
- Methodology
- Read from the tracked record at the last commit.
Limitations
- One run on one evening; no later count exists.
- The key is heuristic: the same sender and subject with no parsed details can merge two different requests; the record does not establish that all 13 were duplicates, and the operator can unmerge.
239 of 239 tests pass at the last commit in the nine reservation and pricing suites, run on 21 Sep 2026 in 173 seconds
Tests passing at the last commit, the nine reservation and pricing suites
verified
- Baseline
- None: the run is the record.
- Sample
- The nine suites (pricing, inbound engine, vendor parser, response guard, reservation service, classify, intent, draft rubric, pipeline runner) at the last commit.
- Methodology
- Run on 21 Sep 2026 from a byte-exact export of the repository at the pin with its own lockfile and configuration, under a guard that makes any network call fail; the summary lines are in the run folder.
Limitations
- Unit tests of pure functions. By name search over the test tree at the pin, no test file references ingestReservationQuotes, autoMergeDuplicates, sendReservationQuote, composeQuoteReply, extractTripFromText or fetchReservationEmails, and the one test file naming a generateDraft imports the outreach service's, not the reservation draft's (row f-test-coverage-search); the status rule, the dedup key, the lifecycle reader, the intent filter and the draft rubric are tested.
- Run locally rather than in the pinned container because the container engine was down that day; the source was proven byte-identical to the pin first.
What happened next
Shipped
- A deterministic pricing engine
- Vendor-form parsing by rules
- Model extraction of free text, always held for a person
- Ready labels at 0.9 and 0.7 for clean priced quotes
- Persistent auto-merge of duplicate forwards
- A cost guard on general-mailbox scanning
Paused
- The 0.90 auto-send threshold, documented as a gate and never invoked
- Automatic customer send
- Computed dead legs, reply stripping, forward merging, round-trip pairing
Unknown
- The model-written reply on the queue path, without a guard
- The part of the rate card a mailbox can reach
- The lookback the cron passes against the one the record describes
Notes on 12 of the 12 items
- A deterministic pricing engine. Documented: yes. Implemented: yes. Deployed: yes. Login-verified: no. Local only: no. calculateQuote is a pure function of typed inputs with every rate a literal in one file; a test asserts that the same inputs return the same quote; the deployed queue priced through it from 20 Jun 2026.
- Vendor-form parsing by rules. Documented: yes. Implemented: yes. Deployed: yes. Login-verified: no. Local only: no. The vendor's notification is recognised by four markers and read by regular expressions, the most complete address kept when a thread repeats them.
- Model extraction of free text, always held for a person. Documented: yes. Implemented: yes. Deployed: yes. Login-verified: no. Local only: no. One model call at temperature zero fills fields and nothing else; any quote built from a model-read field is marked needs review; the June cost episode is the proof it ran on production.
- Ready labels at 0.9 and 0.7 for clean priced quotes. Documented: yes. Implemented: yes. Deployed: yes. Login-verified: no. Local only: no. Ready at 0.9 only for a flat-rate route with no warning that needs a person, 0.7 for any other clean priced quote; the 200-mile minimum warning itself holds a quote.
- The 0.90 auto-send threshold, documented as a gate and never invoked. Documented: yes. Implemented: no. Deployed: no. Login-verified: no. Local only: no. The trust document describes it as enforced; at the last commit it is a constant and a helper that no route, job or service calls.
- Automatic customer send. Documented: yes. Implemented: no. Deployed: no. Login-verified: no. Local only: no. Described as conditional on the gate; at the last commit the only send is a person's click, dry unless a flag is set, and the tracked record never says the flag was set. A plan to start auto-sending high-confidence quotes was set on 1 Jul 2026 in correspondence with the client, for a meeting the next day whose automated summary records no decision; no commit followed.
- Persistent auto-merge of duplicate forwards. Documented: yes. Implemented: yes. Deployed: yes. Login-verified: no. Local only: no. Keyed by reservation number, trip signature or sender and subject; run every ingest cycle and once as a backfill that merged 13 rows into 11 groups.
- A cost guard on general-mailbox scanning. Documented: yes. Implemented: yes. Deployed: yes. Login-verified: no. Local only: no. A 40-term pre-filter and a process-local seen-set before any model call, added in the last commit; the set is cleared on every restart and no post-fix figure exists in the tracked record.
- Computed dead legs, reply stripping, forward merging, round-trip pairing. Documented: yes. Implemented: no. Deployed: no. Login-verified: no. Local only: yes. None of these names occurs at the last commit; they exist only in the builder's uncommitted local batch, never committed or deployed.
- The model-written reply on the queue path, without a guard. The reply a person sends from the queue is drafted by a model and graded by a rubric that passes any dollar figure; the deterministic guard checks only the older inbound flow, and no decision about extending it is recorded.
- The part of the rate card a mailbox can reach. The ingest never sets dead-leg miles, hours, tolls, overnight nights or per-diem days and hardcodes credit-card payment, so most of the rate card is reachable only from the Tester by hand; no decision about it is recorded.
- The lookback the cron passes against the one the record describes. The ten-minute job passes a 24-hour lookback while the service default and the record's cost note say 72 hours; which ran on production in June is not established, and no decision about it is recorded.
Meet the builder
Managing Director and AI Systems Architect
Project contribution
Wrote the pricing engine and the vendor-form parser on 6 May and corrected the rate card five ways after a sync with the client on 21 May; built the Quote Tester on 1 Jun and a deterministic guard over model-written replies on 9 Jun; on 20 Jun put the booking mailbox on a ten-minute ingest with a status per row and wrote the 0.90 gate as a constant nothing calls; on 22 Jun folded every channel into one queue and began scanning two general inboxes; on 23 Jun made duplicate forwards merge persistently; and in the last commit, on 29 Jun, closed the extraction re-loop the bill had revealed with a pre-filter and a seen-set. Computed dead legs, reply stripping and forward merging are later local work, after the last commit, never committed or deployed. The reservation side ran from 6 May to 29 Jun 2026 in two bursts; every commit on its rail carries an AI coding assistant's co-author trailer, a count of trailers, not authorship.
Skills demonstrated
- Where a model may read and where only a function may decideRules and a model fill fields; a pure function prices; a model-read field routes the quote for review (decisions, architecture).
- Explaining a numberLine items with running totals, warnings and a margin band in the Tester and the queue (decisions).
- Bounding autonomyA ready status below the only gate, a gate with no caller, a send that requires a person's action and is dry unless a flag is set (roadmap).
- Stating what was not verifiedThe capability matrix separates implemented from deployed from verified through a login (roadmap, measurement).
From the repository record.
What this project shows
This project shows a Managing Director working as an architect: taking a messy input and drawing the line where a model may read and where only a function may decide. The amount is repeatable for the same inputs; the reading sometimes moved; the queue said which was which. It also shows the discipline of naming what did not ship: the auto-send threshold that nothing invokes, and the model-written reply whose amount the queue path never checks. The other limits are listed once, beside the architecture.
Build one of these
