Training · learner work

The AI Proposes, A Verified Human Decides

A learner built an AI operations platform for SQL Server infrastructure that proposes remediations and cannot apply one. Every write waits for a verified human, and every decision leaves a record that outlives the session that made it.

verified

All student projects
How it works. Narrated by a synthetic voice; the figures it states are the verified metrics recorded below.
Organisation
Colaberry
Industry
Data infrastructure operations
Capability
Governed AI remediation
Status
In progress
Built by
Learner
Published
2026-09-18

The situation

Operations teams do not distrust automation because it is inaccurate. They distrust it because they cannot see what it did, cannot stop it mid-flight, and cannot reconstruct afterwards who allowed it. An assistant that fixes the right thing without leaving a trail is still a thing nobody will let near a database.

That makes the interesting problem a governance one rather than a modelling one. A language model can already read telemetry, correlate failures across services and suggest a remediation. What it cannot do is establish that a person agreed, that the person was who they claimed to be, and that the agreement can be shown to someone else weeks later.

So the design question is not how good the suggestion is. It is where the boundary sits between proposing and acting, whether that boundary is enforced in code or merely written in a policy document, and whether crossing it leaves something durable behind.

What it had to do

  • Put the proposal and the action on opposite sides of a boundary that is enforced by code.
  • Bind an approval to an identity strong enough to be worth binding it to.
  • Persist every decision so it can be reconstructed by someone who was not there.
  • Record the architectural decisions themselves, so the reasoning is inspectable and not just the behaviour.

What constrained it

  • Nothing may write to a monitored system on the strength of a model output alone.
  • An approval must be attributable to a verified identity, not to whoever held a session.
  • The record of a decision has to outlive the session that made it, or it is not an audit trail.
Requirements traceability in the CoreOps Command Center
Requirements traceability in the CoreOps Command CenterThe project's own interface, captured from the repository. Every requirement is traced to the story that fulfils it, with an enforcement state computed from the plan rather than asserted. Among them: the system must recommend actions without executing production changes, and must escalate to a human when its confidence falls below a stated threshold.

Decisions that made the difference

Three choices shaped the system. Each one lives at a specific point in the drawing above.

  1. At ABAC evaluator

    Put the proposal and the action on opposite sides of a boundary enforced by code

    Operations teams distrust automation they cannot see, stop mid-flight or reconstruct afterwards. A model can already read telemetry and suggest a fix; the risk is a write on the strength of a model output alone.

    An ABAC evaluator decides whether an action is permissible under policy at all. A human-in-the-loop queue holds anything that would write until a person acts. A remediation guardrail wraps the action itself.

    Evidence The guardrails directory (architecture narrative); roadmap: approval boundary enforced by an ABAC evaluator and a human-in-the-loop queue, shipped.

    4 of 7guardrail modules carry their own test file, which is the difference between a boundary that is enforced and one that is described. Nothing writes to a monitored system on a model output alone.

  2. At Verified human

    Bind the approval to a verified identity, not to a session

    An approval attributed to whoever held a session cannot answer, weeks later, who agreed. An approval gate is only as strong as the identity standing behind it.

    One decision record commits to multi-factor authentication on the approving account; another binds the approval to a verified identity rather than to a session.

    Evidence ADR-006 and ADR-007 (architecture narrative); roadmap: approvals bound to a verified identity rather than a session, shipped.

    The answer to who approved something outlives the session that produced it.

  3. At Audit trail

    Design the audit trail rather than log it

    A decision that cannot be reconstructed by someone who was not there is not an audit trail, however much was logged.

    Correlate events under a single identifier, persist the trail, and write the architectural decisions themselves down as records beside the code.

    Evidence ADR-002 and ADR-005 (architecture narrative); roadmap: audit trail correlated and persisted, shipped.

    14architecture decision records sit alongside the code: a system whose central claim is that its choices are provable after the fact has written its choices down.

Who built it

  • Learner: built the governance layer, the identity-bound approval and fourteen decision records
The guardrails tab, showing what is and is not enforced
The guardrails tab, showing what is and is not enforcedThe same interface reporting its own limits: of the guardrails shown, one is enforced, one partially, one not yet. The tab labels illustrative rows as sample data rather than presenting them as measured.

The build

  1. Repository created

What was built

The system reads telemetry from SQL Server, SSIS and SSRS, correlates failures across those services, and asks a model to reason about root cause and business impact. That half is unremarkable and is not where the design effort went.

Stack

  • CSS
  • HTML
  • JavaScript
  • Python
  • React
  • Shell
  • TypeScript
  • Vite
More on what was built

The effort went into the guardrails directory. An ABAC evaluator decides whether an action is permissible under policy at all. A human-in-the-loop queue holds anything that would write, and holds it until a person acts. A remediation guardrail wraps the action itself. Four of the seven modules there carry their own test file, which is the difference between a boundary that is enforced and a boundary that is described.

Identity is treated as part of the boundary rather than as a login problem. One decision record commits to multi-factor authentication on the approving account; another binds an approval to a verified identity rather than to a session, so the answer to "who approved this" survives the session that produced it. An approval gate is only as strong as the identity standing behind it.

The audit trail is designed rather than logged. Separate decision records cover correlating events under a single identifier and persisting the trail, which together are what let a decision be reconstructed later by someone who was not present when it was made.

Fourteen architecture decision records sit alongside the code. That is unusual and it is the most telling artefact here: a system whose central claim is that its choices are provable after the fact has written its choices down.

Capabilities

  • Claude code config
  • LLM SDK
  • MCP SDK
  • MCP surface
  • Prompt library

Integrations

  • SQL server
  • Ssis
  • Ssrs
  • MCP tool gateway

Data stores

  • PostgreSQL
  • SQL server
Architecture diagram for The AI Proposes, A Verified Human Decides
A diagram the delivery team drew.

The measurement

This record carries no figures, and that is deliberate.

The governance layer exists and is tested, and the reasoning behind it was written down as it was made. Counting the pieces - the decision records, the guardrail modules, the commits - would say only that, at more length.

What no count here could establish is how well it performs: that the remediations are correct, that an operator adopted it, or that anything downstream improved. Those need a measurement taken in use, which this record does not have and does not pretend to.

What happened next

Shipped

  • Approval boundary enforced by an ABAC evaluator and a human-in-the-loop queue
  • Approvals bound to a verified identity rather than a session
  • Audit trail correlated and persisted

In progress

  • Evidence grounding and structured claim verification

Not pursued

  • Measuring whether the remediations are correct
Notes on 5 of the 5 items
  • Approval boundary enforced by an ABAC evaluator and a human-in-the-loop queue. Both modules carry their own tests. A proposed action is held until a person acts on it.
  • Approvals bound to a verified identity rather than a session. ADR-006 commits to TOTP MFA and ADR-007 binds the approval to identity, so the answer to who approved something outlives the session that produced it.
  • Audit trail correlated and persisted. ADR-002 unifies events under one correlation identifier; ADR-005 commits to persisting the trail.
  • Evidence grounding and structured claim verification. ADR-008 and ADR-009 record the decisions. The records exist; this case study does not establish how much of either is built.
  • Measuring whether the remediations are correct. Nothing in the repository measures the quality of a proposed remediation or whether an operator accepted it. Reporting a figure would mean inventing one.

Meet the builder

Learner and builder

Project contribution

Built the governance layer of an AI operations platform for SQL Server infrastructure: an ABAC evaluator that decides whether an action is permissible, a human-in-the-loop queue that holds any write until a person acts, and a remediation guardrail around the action itself, with four of the seven guardrail modules carrying their own tests; wrote fourteen architecture decision records alongside the code.

Skills demonstrated

  • Boundary designProposal and action on opposite sides of a boundary enforced in code: the ABAC evaluator and the human-in-the-loop queue, both tested (architecture narrative).
  • Identity-bound approvalADR-006 commits to TOTP MFA on the approving account; ADR-007 binds the approval to a verified identity rather than a session.
  • Audit trail designADR-002 unifies events under one correlation identifier; ADR-005 commits to persisting the trail.
  • Decision recordsFourteen architecture decision records sit beside the code, so the reasoning is inspectable and not just the behaviour.
  • Honest measurementThe record carries no figures and says why: nothing in the repository measures whether a remediation is correct or adopted.

From the repository record.

What this project shows

This record shows a governance boundary built as code rather than described in a policy: proposals on one side, writes on the other, an identity-bound approval between them, and fourteen decision records that make the reasoning inspectable. It carries no figures, deliberately. Whether the remediations are correct, whether an operator adopted it and whether anything downstream improved would need a measurement taken in use, which this record does not have and does not pretend to.

Build one of these

Start the program that produced this work.

See the program