Lab

A working matrix of 12 product ideas I've scoped — each one a hypothetical build that maps a real PM skill to a real technical challenge. Most will stay ideas. A few will become projects. All of them are me thinking in public about what's worth making.

These are unbuilt ideas. Built and tested work lives in Projects.

The matrix, sorted by difficulty (hardest first)

Showing 12 of 12 ideas.

Aviation Ops Anomaly FeedExpert

Aviation ops teams drown in green-dashboard fatigue. A feed that only surfaces anomalies (delay drivers, crew-conflict clusters, gate-change cascades) with root-cause breadcrumbs.

Domain modelling, alert-fatigue management

Eval Harness as a ServiceHard

Small AI-PM teams keep rebuilding the same golden-set → re-score → diff pipeline. Offer it as a hosted tool with a drop-in SDK, shadow-traffic mode, and per-prompt-version accuracy deltas.

Developer-tool PM, eval design

PRD → Eval Set ConverterHard

PMs write PRDs; engineers ship without measurable acceptance. Parse a PRD's success criteria into a structured eval set (golden cases, scoring rubric, pass/fail gate) before a single line of code.

Structured thinking, acceptance criteria

Open-Source Contribution ScoutHard

First-time contributors can't find tractable issues. Given a GitHub profile + stack, surface good-first-issues across repos filtered by real complexity (not just the label).

Funnel design, trust metrics

Cost-Per-Answer CalculatorMedium

Every AI team underestimates unit economics until an invoice lands. A live calculator that takes model, tokens-in, tokens-out, cache hit rate, and batch size → returns $/request + $/user/month + envelope check.

Unit-economics thinking, pricing

Churn-Risk ExplainerMedium

Churn models output a score; CS teams need a reason. Produce a one-sentence, feature-level explanation per at-risk user, ranked by recoverability.

Data-driven retention, narrative from numbers

Plant-Health Citizen DatasetMedium

Aarchid's golden set is small (about 200 samples, as reported by the team). Crowdsource a labelled dataset (10k+ photos across failure modes) with contributor attribution and licence clarity.

Marketplace liquidity, contributor incentives

Meeting → Decision LogMedium

Transcripts are noise. Extract decisions made, blockers raised, and owners assigned into a searchable log that survives the meeting.

Information architecture, decision hygiene

Reading-Log → Idea GraphMedium

Highlights sit in Readwise; ideas sit in notes; neither talk. Build a graph view where highlights cluster by concept and surface adjacency between books.

Information design, spaced repetition

AI-PM Interview Prep KitMedium

AI-PM interview loops are new enough that nobody has a clean rubric. Ship a structured practice kit: problem framings, eval-design exercises, cost-envelope drills, mock transcripts with feedback.

Pedagogy, interview-loop design

Interview-Loop Feedback AggregatorEasy

Hiring loops collect 4 independent writeups, then drown in a debrief Slack thread. Aggregate them into a structured rubric summary + dissenting-opinion highlight before the debrief.

Hiring process design, signal vs noise

Micro-SaaS Unit-Economics SimEasy

Solo founders need a back-of-envelope that accepts MRR curve, CAC, churn, infra cost/user, and returns month-by-month cash + break-even date without a spreadsheet.

Business modelling, scenario planning

Source data: ideas.json. Want to build one of these together? Let's talk.