Self-improving AI products.And the teams that build them.
First we agentified our own daily work. Then we baked AI into the products and put it in the hands of users. Now we can build the feedback loop into the product itself: every user action, evaluation result and outcome is traced back to the prompts, skills, thresholds and weights that shaped the result, and the next version takes shape while we sleep.
Each AI product below lets its signals flow back through the engine and change it. This is the pillar I would bring to your team.
The forward pass is what every product has: retrieve, prompt, generate, apply the rules. The backward pass is what makes it self-improving: each signal is verified, joined to the result that caused it, and then changes one layer, which ships as a new version. The model’s self-reflection takes a separate path, into a roadmap that people curate.
Let the product be its own product manager for what can be measured, and free the people for what has to be judged. Every product built this way carries a flywheel: a versioned unit of output, a signal from outside the model, attribution back to whatever produced the output, a score for each source that shaped it, and a change to one layer that ships as a new version.
The page behind each product shows how its loop is built: what is captured, how it is attributed to the thing that produced it, and how a change ships.
Four of the products below run this loop. The pattern page shows how, one part at a time.
Commits to my own products, updated daily. Client work is not included. Hover a square to see which products each day went to.
Tales from the loop
Work
All in production since 2024, each designed, built and run end to end with Claude Code. Four of them run the loop.
ProductWhat it doesBuilt withStatus
SuperAgentiChat superagentichat.comFinancial research agentA financial research agent that turns a question into a professional-grade analysis report. An agentic loop with dynamic context management works across 40+ tools, numerous data sources and 70+ built-in skills. It reads your files, runs code, charts the data and writes output files, with proprietary OSINT analysis and bias detection built in.Role, features, how it learnsHide detailsVercel AI SDK, Next.js, Neon Postgres, Paddle billingLive
My role
Everything from the research loop to the checkout flow.
Professional-grade analysis reports with financial data visualization
Works on your own files, executes code and produces output files
Custom user skills on top of the built-in ones
Full billing lifecycle: checkout, portal, webhooks, cancellation
Learns from use. Every answer is judged and voted on, and the winners become the next version of the prompts and tools.
DropRenew droprenew.comDecision systemAn operating system for a domain portfolio. A weekly AI renewal committee weighs the evidence on every expiring name and shows its reasoning before you renew or drop. Around it: an AI wizard that sorts the portfolio into investment themes, a weekly market-intelligence issue matched to those themes, pricing strategy tools, and a post-mortem on every name you let go.Role, features, how it learnsHide detailsTypeScript, Supabase, Replicate, Inngest jobs, registrar APIsLive
My role
What started as a personal project to manage my own domain portfolio grew into a consumer product. A scouting mechanism harvests publicly shared knowledge from the domain-investing world and distils it into market consensus and opinions that feed the roadmap.
Portfolio tracking across registrars: costs, valuations, traffic and listings in one table
Weekly market intelligence: comparable sales, trademark radar and emerging vocabulary, matched to your themes
Post-mortem on drops: who re-registered the name, what they built, and whether the call was right
Learns from use. Every override, and every real sale or drop, is scored against the engine that made the call. A change ships only if it beats it.
OPERATIVYOperations engineA marketing engine for an owner-operator running several brands. It feeds on market news and a corpus of marketing knowledge to produce content for multiple channels, aimed at the specific goals of each brand. The work happens overnight, and the operator's morning is an hour or two of yes, no or edit calls on prepared drafts.Role, features, how it learnsHide detailsRemix, TypeScript, Neon Postgres, Graphile Worker, RailwayInternal
My role
The tool I use to run my own brands, so I am its first and most demanding operator. Every morning's yes, no or edit is mine, and so is the cost of a weak post.
Overnight pipeline: news ingest, idea synthesis, drafts and reconciliation
Morning triage across every brand, each with its own thesis and voice
A browser helper fills the LinkedIn composer with an approved draft, and you click Post
Learns from use. Two loops. Every post carries a tracking link, so what it brings back is traced to the plays that shaped it, and plays that lose more than they win are retired. And the engine drafts more candidates than get published, so each clearance teaches it the operator's preferences.
SPINES getspines.comBook discoveryA new kind of book discovery for serious readers. Point a camera at a bookshelf and every spine becomes a known book. SPINES then shows which books keep appearing together on real shelves, and leads you to the one at the edge of what you know instead of the most popular one. Shelfology, its analysis layer, reads your own shelf back to you: what you collect, which spines are rare, and whose library resembles yours. Other agents can read shelves too, through a public API, an SDK and an MCP server.Role, features, how it learnsHide detailsTypeScript, Roboflow, vision models, MCP, Node SDKLive
My role
A passion project, born of being an avid non-fiction reader. I am building a first-of-its-kind database of how people's reading choices are reflected in their home libraries, one shelf at a time.
Web app, SDK and MCP server from one codebase
Curated persona and pairing engine on top of recognition
Learns from use. Today the text on each spine is read to identify the book. Every approved shelf turns those readings into labelled spine images, the training set for tomorrow's vision classifier that knows a book by its spine alone. The shared catalogue grows with each shelf too, so the next photo matches faster.
Self Driving Cars 101 selfdrivingcars101.comPhysical AI communityA global, open community and network of meetup groups for everything about autonomous vehicles: the technology, the business, the debate and the impact. Volunteer-run since 2016, 6,000+ members, and no paid sponsors steering the agenda. The website is the community's hub: its resources, learning material, weekly newsletter and a locator for the ten local meetup groups.Role, features, how it learnsHide detailsNext.js, Postgres, Resend, PlaywrightLive, growing
My role
Founder and general manager. I started it in 2016 as a home for the people building physical AI on the road, and I still run it.
Weekly Briefing with double opt-in, an archive, and an admin studio to assemble each issue
Learn hub: thirteen lessons and a glossary, with a generated llms.txt so agents can read them too
Every item below lives in the code of the products above.
Something old
Good engineering
The discipline I asked of engineering teams, applied to myself and to the agents I direct.
Decisions before code. Architecture decision records in every AI product, and a hard gate: design, approval, plan, then code. docs/adr/
Tests first. Model providers are mocked, so a feature fails loudly without spending a token. test-utils/mocks/llm
Every lesson becomes a check. Repeated incidents turn into lint rules that block the commit. Made mechanical, not memorised. no-unoptimized-cover-img
Known issues, written down. Living known-issues and lessons-learned files, with IDs, in every AI product. KNOWN_ISSUES.md
Jobs that watch themselves. Scheduled health checks, backups and alert digests that stay silent when all is well. alert_sentinel
A dated reason to stop. Kill criteria with a number and a date, written into the spec before the build starts.
Something new
Managing intelligence
Models are powerful and unreliable. The system around them decides which one you get.
The model never does the math. Calculations run in code. The model scores or referees only where judgment is needed. forces.ts
Deterministic judges first. Code checks such as the balance-sheet identity run beside the LLM judges, and the judges are calibrated against real outcomes. code-checks.ts
Evals on a golden set. A prompt version ships only after a paired A/B run against curated real cases. eval_runs
Every prompt versioned, every call traced. Prompt ID and version stamped on each call, with the full request trace kept. prompt-registry.ts
Cost is a metric. Budgets, a cost stamp on every call, and a weekly check for provider price drift. pricing-drift-checker.ts
Models earn their place. Candidate models are benchmarked against a hand-built reference before routing allows them in. benchmark-dense-shelf.ts
Something borrowed
Learning from others
Most hard problems were half-solved somewhere else first. Find where before building.
Research before design. A written research note comes before each major design decision. docs/research/
Scouting sprints. Outside agents from several labs get the same brief, without seeing the system's internals, and hunt for tools, skills and methods. Every proposal is scored and deduplicated, and every agent is graded on a scorecard. ResourceScout
Best of breed. For a chosen capability, several agents research it in depth. The design takes the best formula, threshold and code from each, and one convergence pass decides what ships. playscout/CONVERGENCE.md
Outside agents as reviewers. The same design question goes to several models from other labs, and their answers are merged into one synthesis. adr-003-synthesis.md
Proofs of concept before commitments. A scripted spike settles the choice between approaches. poc-two-pass.ts
Reviews with receipts. Multi-agent code reviews where every finding gets a severity, a status, and a written reason if it is not fixed. CODE_REVIEW
Something blue
AI security
Anything the model reads can be an attack. Defend like a blue team.
A written threat model. Prompt injection is designed for up front, not patched after. ADR-011
User space, walled off. Input is schema-checked and size-limited at the boundary. Everything a user writes is then fenced as data, below the system's own rules. The fence holds because user text cannot contain it, and only blocks carrying a per-conversation secret count as system instructions. ugc-fence.ts
<user_generated_content type="personal_profile" trust_level="untrusted">
...anything the user wrote...
</user_generated_content>
Untrusted content is neutralised. Web pages and outside data are cleaned before they reach a prompt. untrusted.ts
Guarded output. Responses are checked for data exfiltration and system-prompt leaks before they leave. output-guard.ts
Abuse limits. Rate limits and automatic blocking of misbehaving API keys. api-abuse.ts
Hostile tests. Tenant isolation is attacked by dedicated fixtures in CI. hostile-fixtures
Also writing inGet Spines, the SPINES publication on reading and thinking, and Self Driving Cars 101, the community’s publication on autonomy.
Four ways in
Work with me
For founders, product teams and anyone with an AI product that needs to exist.
Self-improving AI productsDesign the loop that turns usage into a better product: what to capture, how to attribute it, and where a human gates the change. See the pattern.
Zero to liveAn idea becomes a deployed product: scope, architecture decisions, build, tests, billing, launch. Weeks, not quarters.
Fractional product leadOwn discovery and the roadmap for an early-stage team a few days a week, or step in to lead product when the problem is bigger than that.
Claude Code for product teamsSet your team up to work the way I do: decision records, test-first, agent workflows, custom lint rules. Hands-on, not slides.
Built by one person, twenty-two years in: chips, then hardware, then devices, then AI software, then agents. About me
Build the hard part.
Send a paragraph about what you're building and where you are with it. I'll reply within a working day with whether I can help, a rough shape, and what I'd want to know first.