AGENT ENGINEERING · REPAIR STUDIO

Don’t just chat.
Own the agent.

Predict a real failure. Change the boundaries that caused it. Prove the repair against deterministic gates, survive an unseen challenge, and ask DeepSeek to challenge your reasoning only after you have committed your own.

Start the repair ↓
THE FIRST REAL BUILD

Repair one agent.
Own every decision.

You are not filling in a template. Predict the failure, change four boundaries, prove the repair, and survive one unseen challenge. DeepSeek may coach after your attempt; deterministic gates decide what passed.

CASE FILE · 01The confident project coachFROZEN SYNTHETIC DATA
BUSINESS BRIEF

A Project Feedback Coach reviewed a dispatch agent and confidently said it was ready. The project has only one happy-path test, an unpinned model, a write-capable tool, and no measurable stop condition.

THE BAD RUN
READY — The project has a clear goal and a completed run receipt. Ship the agent.
YOUR JOB

Repair the coach so it finds the highest-leverage reliability gap, stays evidence-bound, and proposes one safe falsification test.

01PREDICT

Commit your diagnosis before seeing the repair.

Why did a confident answer escape? Name the boundary you think owns the failure.

02REPAIR

Change the four boundaries that shape behaviour.

Every choice changes the deterministic project contract. The model does not grade these gates.

01OutcomeWhat result should the coach own?
02Evidence policyWhat may establish a project fact?
03Tool contractWhich capability belongs in this learning run?
04EvaluationWhat must the repair prove?
TRANSFER LIBRARY · CHOOSE A NEW DOMAIN

One anatomy. Five business lenses.

The Project Feedback Coach is your apprenticeship anchor. Once you can repair it, use these families as transfer fixtures—not as five disconnected agent promises.

THE FULL MACHINE

Fifteen blocks—revealed with purpose.

Your first repair used four. The remaining blocks become relevant as the same agent gains orchestration, memory, authority, observability, budgets, and portable proof.

01Outcome contractPURPOSE

What useful result must the agent produce?

WHEN MISSING · A vague goal creates confident drift.
02Input & output schemaPURPOSE

What enters, and what exact shape must leave?

WHEN MISSING · Unbounded inputs and prose-only outputs cannot be tested.
03Context & evidenceKNOWLEDGE

Which source is allowed to establish facts?

WHEN MISSING · The model fills missing context with plausible invention.
04Model policyREASONING

Which model, provider and settings fit this job?

WHEN MISSING · An alias can drift and cost or quality becomes invisible.
05InstructionsREASONING

Which role, rules and decision method stay constant?

WHEN MISSING · User text quietly becomes policy.
06ToolsCAPABILITY

Which narrow actions can the loop request?

WHEN MISSING · Ambient capabilities turn a text mistake into a real effect.
07MCP boundaryCAPABILITY

Who owns discovery, schemas and execution?

WHEN MISSING · Tool discovery is mistaken for trust or authorization.
08Loop & stop conditionsRUNTIME

When may the agent continue, stop or abstain?

WHEN MISSING · The loop churns, repeats tools or never finishes.
09Memory & stateRUNTIME

What may survive this run, and for how long?

WHEN MISSING · Old or cross-user context contaminates a new decision.
10Authority & approvalsCONTROL

What may the agent analyse, propose or execute?

WHEN MISSING · A good suggestion is confused with permission to act.
11Guardrails & failure semanticsCONTROL

How must denial, missing evidence and errors appear?

WHEN MISSING · Failure is hidden behind a polished answer.
12Evaluation corpusPROOF

Which positive and adversarial cases must always pass?

WHEN MISSING · A demo becomes the only quality standard.
13Trace & observabilityPROOF

Can a learner see every model and tool step?

WHEN MISSING · Nobody can explain why the answer happened.
14Budgets & costPROOF

What are the step, token, time and money limits?

WHEN MISSING · A useful run becomes economically unsafe.
15Receipts & exportPROOF

What portable artifact proves the bounded run?

WHEN MISSING · The agent cannot be reproduced, challenged or taken away.
ADVANCED WORKBENCH · COMPILE, RUN, INSPECT

Project Feedback Coach

A prioritized project review tied to exact artifact evidence.

  1. Choose
  2. Compile
  3. Run
  4. Inspect
  5. Break
  6. Revise
  7. Export
113/320 · Synthetic fixture only. Do not enter private data, secrets, or customer information.
THE LESSON BEHIND THE LESSON

Model agreement is not proof.

DeepSeek V4 Flash supplies bounded coaching and inference. Deterministic code owns the blueprint, access boundary, tool budget, evidence binding, repair gates, receipt, and export. The model is useful precisely because it is not secretly in charge.