Lab
A working matrix of 12 product ideas I've scoped — each one a hypothetical build that maps a real PM skill to a real technical challenge. Most will stay ideas. A few will become projects. All of them are me thinking in public about what's worth making.
These are unbuilt ideas. Built and tested work lives in Projects.
The matrix, sorted by difficulty (hardest first)
Showing 12 of 12 ideas.
Aviation Ops Anomaly FeedExpert
Aviation ops teams drown in green-dashboard fatigue. A feed that only surfaces anomalies (delay drivers, crew-conflict clusters, gate-change cascades) with root-cause breadcrumbs.
Domain modelling, alert-fatigue managementEval Harness as a ServiceHard
Small AI-PM teams keep rebuilding the same golden-set → re-score → diff pipeline. Offer it as a hosted tool with a drop-in SDK, shadow-traffic mode, and per-prompt-version accuracy deltas.
Developer-tool PM, eval designPRD → Eval Set ConverterHard
PMs write PRDs; engineers ship without measurable acceptance. Parse a PRD's success criteria into a structured eval set (golden cases, scoring rubric, pass/fail gate) before a single line of code.
Structured thinking, acceptance criteriaOpen-Source Contribution ScoutHard
First-time contributors can't find tractable issues. Given a GitHub profile + stack, surface good-first-issues across repos filtered by real complexity (not just the label).
Funnel design, trust metricsCost-Per-Answer CalculatorMedium
Every AI team underestimates unit economics until an invoice lands. A live calculator that takes model, tokens-in, tokens-out, cache hit rate, and batch size → returns $/request + $/user/month + envelope check.
Unit-economics thinking, pricingChurn-Risk ExplainerMedium
Churn models output a score; CS teams need a reason. Produce a one-sentence, feature-level explanation per at-risk user, ranked by recoverability.
Data-driven retention, narrative from numbersPlant-Health Citizen DatasetMedium
Aarchid's golden set is small (about 200 samples, as reported by the team). Crowdsource a labelled dataset (10k+ photos across failure modes) with contributor attribution and licence clarity.
Marketplace liquidity, contributor incentivesMeeting → Decision LogMedium
Transcripts are noise. Extract decisions made, blockers raised, and owners assigned into a searchable log that survives the meeting.
Information architecture, decision hygieneReading-Log → Idea GraphMedium
Highlights sit in Readwise; ideas sit in notes; neither talk. Build a graph view where highlights cluster by concept and surface adjacency between books.
Information design, spaced repetitionAI-PM Interview Prep KitMedium
AI-PM interview loops are new enough that nobody has a clean rubric. Ship a structured practice kit: problem framings, eval-design exercises, cost-envelope drills, mock transcripts with feedback.
Pedagogy, interview-loop designInterview-Loop Feedback AggregatorEasy
Hiring loops collect 4 independent writeups, then drown in a debrief Slack thread. Aggregate them into a structured rubric summary + dissenting-opinion highlight before the debrief.
Hiring process design, signal vs noiseMicro-SaaS Unit-Economics SimEasy
Solo founders need a back-of-envelope that accepts MRR curve, CAC, churn, infra cost/user, and returns month-by-month cash + break-even date without a spreadsheet.
Business modelling, scenario planningSource data: ideas.json. Want to build one of these together? Let's talk.