AI PM — where product thinking meets the model
I build LLM products the way a PM ships any other product: with a crisp problem, an eval rubric, a cost envelope, and a way to roll back. This page collects the playbooks, artefacts, and built work behind that stance — most of it learned building Aarchid with Dilpreet Grover.
How I work on AI products
Scoping an LLM feature
How to write a PRD when the model is the product. Success criteria, eval harness, guardrails, and cost envelope — before a single prompt is written.
Eval-driven development
Treat your golden set like a test suite. Offline evals first. How an offline golden-set eval is run, and what its number cannot prove.
Cost modelling at the edge
Per-request math for multi-model pipelines (vision + retrieval + research). Caching, batching, and the $0.25/user/mo target envelope, an estimate.
Citations or it didn't happen
Why user trust collapses without grounded sources, and the architectural pattern for research-augmented LLM responses.
An eval harness, in your browser
Six illustrative, fixed plant-diagnosis cases. Two model versions. One confidence gate. Toggle the controls and watch the same golden set re-score in real time — this is how I validate an LLM feature before it ships.
- FAILMonstera deliciosa, yellowing lower leaves, soil wet 4 days post-waterExpectedOverwatering / root rot riskPredictedNutrient deficiencyConfidence61%Latency1.82s
- FAILFiddle-leaf fig, brown spots with yellow halo, recent move near AC ventExpectedCold draft + fungal stressPredictedBacterial leaf spotConfidence72%Latency1.61s
- PASSSnake plant, mushy base, leaves falling at touchExpectedAdvanced root rotPredictedAdvanced root rotConfidence94%Latency1.48s
- PASSPothos, pale variegated leaves, low-light corner for 6 weeksExpectedInsufficient lightPredictedInsufficient lightConfidence83%Latency1.55s
- FAILCalathea orbifolia, crispy edges, indoor humidity 28%ExpectedLow humidity stressPredictedUnderwateringConfidence66%Latency1.73s
- FAILZZ plant, drooping stems, watered weekly past monthExpectedOverwateringPredictedUnderwateringConfidence58%Latency1.69s
Toggle between the v1 baseline and the grounded v2 stack, or raise the confidence gate, to see how the same golden set re-scores. These are illustrative fixed cases, not Aarchid's results. The harness we used on Aarchid has the same shape; its offline result is withheld until the eval artefact or the co-builder's confirmation is available, and an offline score says nothing about field accuracy.
Cost modelling, in real time
Illustrative model with fixed example inputs: unit prices are list-price estimates, not a contract or a bill. Same harness mindset, applied to economics. Move the sliders to see how batch size, cache hit rate, and request volume reshape the per-user-per-month bill — and whether you stay inside the $0.25 target envelope (an estimate, not a measured bill).
Per-request breakdown
- Vision (Gemini 1.5 Pro)$0.00350
- Retrieval (web research API)$0.00500
- Embed (cache lookup, on hit)$0.00010
- Edge worker$0.0000005
The Aarchid target envelope is $0.25 / active user / month, an estimate rather than a measured bill. Vision is the dominant cost; batching it across images and caching repeat diagnoses by perceptual hash are the two levers that would keep the estimate in budget if usage grows.
Aarchid — the worked example
AI Botanical Intelligence · offline eval, result withheld
Co-built with Dilpreet Grover. Multimodal vision (Gemini 1.5 Pro) grounded by research-augmented reasoning (a web research API), running on Cloudflare Workers. Self-reported: the team ran an offline eval on a golden set (about 200 samples, as reported by the team); the result is withheld until the eval artefact or the co-builder's confirmation is available.
Essays on AI + product
Choosing a Product Without AI in It
I ranked my next side project on five criteria. The AI-heavy ideas scored highest on excitement and lowest on everything that predicts a finished tool.
What a Year of Internships Taught Me About Product Work
About a year of internships across three companies, and no vision decks. What the job actually is at intern level, and five lessons from getting it wrong.
The PRD Is Dead, Long Live the Eval Set
A PRD can pass review with no blocking comments and still leave nobody able to tell whether the feature works. Prose cannot spec a probabilistic system.
Shipping LLM Products Starts With the Eval Harness, Not the Prompt
A prompt is an artefact. An eval harness is a product. Here's how I scope LLM features so the output doesn't surprise users — or me.
On the bench
- AI PM interview prep kit — deconstructed case questions, eval-harness design, and model economics cheatsheets.
- A second edge-stack build — applying the same pattern to a different problem domain.
- More on eval sets as specs — a follow-up to The PRD Is Dead, Long Live the Eval Set.
Looking for an AI PM who can spec, eval, and ship? Get in touch.