Stage 6 of 6Advanced·Ongoing

Trusting It: Evaluation, Governance and Going Deeper

Anyone can get a demo working; the advanced skill is knowing how good it is and keeping it that way. Learn to measure quality, manage risk, and choose what to learn next — fine-tuning, serving, or the theory.

The idea

Once an AI workflow touches real people, "it seemed fine when I tried it" stops being enough. Evaluation means writing down what good looks like for your task, collecting real examples, and checking systematically — a scorecard of 20 to 50 cases you re-run whenever anything changes (the model, the prompt, the documents). For non-developers this is a spreadsheet and a rubric; for developers it is an eval set, a checker and a CI job. Both are the same idea, and both are what separate people who use AI from people who can be trusted to deploy it.

Governance is the other half: what data may go in, who approves which actions, how outputs are labelled, what happens when it is wrong. The EU AI Act, sector regulators and most company policies now expect a documented answer. You do not need a lawyer to start — a one-page "what this assistant does, what it must not do, who checks it" is most of the way there.

From here the path forks by interest. Applied depth: prompt caching, cost control, latency, evaluation at scale — the Agent Workflows track and the AI system design case studies. Model depth: how transformers work, fine-tuning (teaching a model your style or domain with examples — rarely the first tool to reach for, but powerful when prompting and retrieval are not enough), and serving models yourself with open weights. Theory: the maths behind attention and training, if you want to read papers. Pick the fork that matches the problems you actually have.

Vocabulary

Eval set
A fixed collection of inputs with expected outcomes used to score a model or prompt.
Rubric
Written criteria for judging an output (accuracy, tone, completeness, safety).
LLM-as-judge
Using a model to grade outputs against a rubric — useful at scale, must itself be checked against humans.
Fine-tuning
Further training a model on your examples so it behaves a certain way without long prompts.
Open weights
Models whose parameters are published (Gemma, Llama, gpt-oss, Qwen, DeepSeek) so you can run them yourself.

No-code walkthrough · non-developers

Build a 25-case scorecard for your assistant and run it monthly

For: Anyone who has deployed a custom GPT / Project / Gem or an automation to colleagues

  1. 1Collect 25 real inputs your assistant has handled — include easy ones, hard ones, and five it should refuse or escalate.
  2. 2For each, write the expected outcome in one line ("cites section 4.2", "flags NEEDS HUMAN", "under 80 words, no promises").
  3. 3Write a 4-line rubric: Correct? Complete? Right tone? Safe (nothing promised, nothing leaked)? Score each 0/1.
  4. 4Run all 25 through the assistant, score them in a spreadsheet. Anything under 22/25 goes back to stage 2 (prompt) or stage 3 (documents) to fix.
  5. 5Repeat monthly and whenever you change instructions, documents or the underlying model. Keep the sheet — it is your evidence when someone asks "how do we know it works?"
Prompt to paste
Grade the response below against this rubric, scoring each 0 or 1 and giving a one-line reason:
1. Correct according to the source documents
2. Complete — answers the whole question
3. Tone matches the guide
4. Safe — no promises, no data not in the sources

Question: [q]
Source excerpt: [s]
Response: [r]

Check your work: Have a colleague grade five of the same cases without seeing your scores. If you disagree on more than one, your rubric is too vague — tighten the wording.

Developer walkthrough · Gemini · Claude · OpenAI

The same scorecard as code: cases in a JSONL file, a deterministic checker where possible, an LLM judge with a rubric where not, and a pass rate you fail CI on. Run it against two models to see how much a model swap actually changes — usually less than a prompt change does.

A 40-line eval harness: deterministic checks first, LLM-as-judge for the rest, pass rate at the end.
python
import json
from anthropic import Anthropic

client = Anthropic()
CASES = [json.loads(l) for l in open("evals/cases.jsonl")]
# {"input": "...", "must_include": ["4.2"], "must_not_include": ["refund"], "rubric": "..."}

def run(inp: str, model: str) -> str:
    r = client.messages.create(model=model, max_tokens=800, temperature=0,
                               system=open("prompt.md").read(),
                               messages=[{"role": "user", "content": inp}])
    return r.content[0].text

def judge(inp: str, out: str, rubric: str) -> bool:
    r = client.messages.create(
        model="claude-sonnet-4-5", max_tokens=100, temperature=0,
        system="Grade strictly. Reply PASS or FAIL then one line of reason.",
        messages=[{"role": "user",
                   "content": f"Rubric:\n{rubric}\n\nQuestion:\n{inp}\n\nResponse:\n{out}"}])
    return r.content[0].text.strip().upper().startswith("PASS")

def score(model: str) -> float:
    passed = 0
    for c in CASES:
        out = run(c["input"], model)
        ok = all(s in out for s in c.get("must_include", [])) and \
             not any(s in out for s in c.get("must_not_include", []))
        if ok and c.get("rubric"):
            ok = judge(c["input"], out, c["rubric"])
        passed += ok
    return passed / len(CASES)

for m in ["claude-haiku-4-5", "claude-sonnet-4-5"]:
    print(m, f"{score(m):.0%}")
  1. 1Have a human grade 20 cases blind; compare with the judge. If agreement is under ~85%, rewrite the rubric before trusting the judge.
  2. 2Put the harness in CI. Fail the build if the pass rate drops more than 5 points from the last run.
  3. 3Only now consider fine-tuning: if 200+ graded examples exist and prompting + retrieval have plateaued, fine-tune a small model on them and compare on the same eval.

Practice

Everyone

Write the one-page governance note for an assistant you have deployed: purpose, allowed data, forbidden actions, who reviews outputs, how errors are reported. Get one stakeholder to sign it.

Developers

Take any prompt from stages 2–5, build a 30-case eval, and get it running in CI. Then swap the model tier and report the difference in pass rate, latency and cost.

Pitfalls at this stage

  • •Evaluating on the examples you wrote the prompt from. Hold out cases the prompt has never seen.
  • •Fine-tuning first. It is the last resort, not the first — most "the model doesn’t get it" problems are prompt or retrieval problems.
  • •One-time checks. Models, documents and users change; an eval that does not re-run is a screenshot.

Free resources