AI PM — where product thinking meets the model

I build LLM products the way a PM ships any other product: with a crisp problem, an eval rubric, a cost envelope, and a way to roll back. This page collects the playbooks, artefacts, and built work behind that stance — most of it learned building Aarchid with Dilpreet Grover.

An eval harness, in your browser

Six illustrative, fixed plant-diagnosis cases. Two model versions. One confidence gate. Toggle the controls and watch the same golden set re-score in real time — this is how I validate an LLM feature before it ships.

33%Accuracy, illustrative set (2/6)
72%Avg confidence
1.65sAvg latency
  • FAILMonstera deliciosa, yellowing lower leaves, soil wet 4 days post-water
    ExpectedOverwatering / root rot risk
    PredictedNutrient deficiency
    Confidence61%
    Latency1.82s
  • FAILFiddle-leaf fig, brown spots with yellow halo, recent move near AC vent
    ExpectedCold draft + fungal stress
    PredictedBacterial leaf spot
    Confidence72%
    Latency1.61s
  • PASSSnake plant, mushy base, leaves falling at touch
    ExpectedAdvanced root rot
    PredictedAdvanced root rot
    Confidence94%
    Latency1.48s
  • PASSPothos, pale variegated leaves, low-light corner for 6 weeks
    ExpectedInsufficient light
    PredictedInsufficient light
    Confidence83%
    Latency1.55s
  • FAILCalathea orbifolia, crispy edges, indoor humidity 28%
    ExpectedLow humidity stress
    PredictedUnderwatering
    Confidence66%
    Latency1.73s
  • FAILZZ plant, drooping stems, watered weekly past month
    ExpectedOverwatering
    PredictedUnderwatering
    Confidence58%
    Latency1.69s

Toggle between the v1 baseline and the grounded v2 stack, or raise the confidence gate, to see how the same golden set re-scores. These are illustrative fixed cases, not Aarchid's results. The harness we used on Aarchid has the same shape; its offline result is withheld until the eval artefact or the co-builder's confirmation is available, and an offline score says nothing about field accuracy.

Cost modelling, in real time

Illustrative model with fixed example inputs: unit prices are list-price estimates, not a contract or a bill. Same harness mindset, applied to economics. Move the sliders to see how batch size, cache hit rate, and request volume reshape the per-user-per-month bill — and whether you stay inside the $0.25 target envelope (an estimate, not a measured bill).

$0.00472Effective / request
$0.038/ user / month · in envelope
$189Monthly run-rate

Per-request breakdown

  • Vision (Gemini 1.5 Pro)$0.00350
  • Retrieval (web research API)$0.00500
  • Embed (cache lookup, on hit)$0.00010
  • Edge worker$0.0000005

The Aarchid target envelope is $0.25 / active user / month, an estimate rather than a measured bill. Vision is the dominant cost; batching it across images and caching repeat diagnoses by perceptual hash are the two levers that would keep the estimate in budget if usage grows.

Aarchid — the worked example

AI Botanical Intelligence · offline eval, result withheld

Co-built with Dilpreet Grover. Multimodal vision (Gemini 1.5 Pro) grounded by research-augmented reasoning (a web research API), running on Cloudflare Workers. Self-reported: the team ran an offline eval on a golden set (about 200 samples, as reported by the team); the result is withheld until the eval artefact or the co-builder's confirmation is available.

Read the case study →

On the bench

  • AI PM interview prep kit — deconstructed case questions, eval-harness design, and model economics cheatsheets.
  • A second edge-stack build — applying the same pattern to a different problem domain.
  • More on eval sets as specs — a follow-up to The PRD Is Dead, Long Live the Eval Set.

Looking for an AI PM who can spec, eval, and ship? Get in touch.