A common mistake
Picture this: a team decides to "add AI". Someone writes a prompt. The demo works on three carefully chosen inputs. The PR ships to 10% of users. Two weeks later, the complaints arrive: the model confabulated a source, gave contradictory advice, and broke gracefully exactly zero of the times it should have.
The bug isn't the prompt. The bug is that there was never an eval.
A prompt is an artefact. An eval harness is a product.
On Aarchid — the AI botanical diagnosis platform I co-created with Dilpreet Grover — the team ran an offline eval on a golden set of about 200 hand-labelled samples (as reported by the team). I am withholding the result until the eval artefact or my co-builder's confirmation is available.
The order I recommend, and the one this post argues for, is to have three things before a single production prompt is written:
- A golden set of hand-labelled inputs with ground truth. For a plant-diagnosis product: species, condition, severity, recommended action.
- A rubric with a few separate dimensions. Aarchid's has four: diagnostic accuracy, citation fidelity, severity calibration, and latency.
- A scoring pipeline that runs the full chain (for Aarchid, vision → research → synthesis) against the golden set and emits a comparable score.
Only then iterate on prompts. Every change — model version, system prompt, retrieval strategy, temperature — runs through the harness, and each release is judged by the score, not by how good the demo looks.
What "good" looks like for an LLM feature PRD
Most PRDs for LLM features read like: "The model will answer questions about X." That's a capability statement, not a spec.
The spec that actually ships has five parts:
1. Behaviour contract
What must the output always contain? What must it never contain? On Aarchid:
- Always: health score (1–100), severity tier, at least one cited action.
- Never: species-identification claims below a confidence floor, recommendations without a source URL, hedging language that obscures severity.
2. Golden set
A small (50–200), diverse, hand-labelled dataset. Add a case each time you find a failure the set misses. Revisit it every release.
3. Rubric
3–6 axes, each scored 0–5 or 0–1. Resist the urge to collapse it into one number too early — the axes teach you where you're weak.
4. Guardrails
What happens when the model falls off a cliff? Retry, degrade, escalate, or refuse. Aarchid's spec says to decline when the vision model's confidence on species ID is low — better silence than a confident wrong answer. That is a design rule in the spec; how often it fires in the field is not measured.
5. Cost envelope
Per-request maths, including retries and retrieval calls. If you can't afford your feature at P95 usage, you don't have a feature.
The eval loop in practice
The loop has four steps:
- Golden set goes in.
- Pipeline runs: vision plus retrieval.
- Rubric scores each output on four axes.
- Score and diff against the baseline, then back to step 1 on the next change.
Every prompt change, every model upgrade, every retrieval tweak runs the loop. Regressions are caught before a user sees them. Improvements are measurable.
This is not novel. Traditional ML teams have done this forever. The mistake is thinking that because LLMs are "just prompts", they don't need the same discipline. They need more, because the failure surface is larger and the confidence is higher.
Three lessons from Aarchid
1. Build the harness before the feature. That is the recommended order, not a claim about the order Aarchid followed. It feels slow. It isn't. Every other decision gets faster.
2. Grow the golden set on purpose. Aarchid's set is about 200 samples, none from production traffic (as reported by the team). Each case you add should cover something the earlier set missed and become a permanent regression check.
3. Cost is a first-class axis. Track it next to accuracy, not after it. Aarchid's cost target was about $0.25 per active user per month, an estimate rather than a measured bill. The offline accuracy result is withheld, as above, and field accuracy is not measured.
What to put on the PR-FAQ
When I scope an LLM feature, I ask for three things before I'll call the spec ready:
- Show me the golden set.
- Show me the rubric.
- Show me the cost-per-request maths with retries included.
If those three don't exist, we don't have a feature — we have a demo that will embarrass us at scale.
The short version
Ship the eval harness. Then ship the feature.
More of this thinking lives on my AI PM page. If you're wrestling with scoping your first LLM feature, get in touch — I like these conversations.
