AI lab

What I build with AI agents, in my own time

Since February 2026 I have used AI coding agents to build software end to end: a governance and routing layer for AI agents, a control plane that connects assistants to infrastructure, a learned routing model, iPhone apps that are live on the App Store, and sites built with and for the people in my life. That is about 7,500 commits in eight months, backed by more than 2,700 test files.

These are personal projects, built outside work and independent of Acronis. My day job is partnerships; this is how I stay close to what AI platforms can actually do.

AI systems

Three systems that work together: one governs and routes what agents do, one gives assistants a safe way to act on infrastructure, and one learns how to spend less on the work.

Feb–Sep 2026

BrainstormRouter

An authorization, evidence and routing layer for AI agents. Every agent action runs under a cryptographically signed, scoped capability grant, and every decision, allowed or denied, lands in a tamper-evident ledger. Inside that boundary the router picks the best model for each turn.

  • Ed25519-signed capability grants with verifiable delegation: an agent can pass a narrower slice of its authority to another agent, never a wider one.
  • Hash-chained, append-only audit ledger, so denials are evidence too.
  • Thompson-sampling model routing across 55 curated models from nine providers, with real-time budgets and memory refresh.
  • Drop-in OpenAI- and Anthropic-compatible API, plus a Model Context Protocol (MCP) server.
How BrainstormRouter handles an agent request An agent request first meets a signed capability grant check. Allowed requests go to the router, which uses Thompson sampling to choose among 55 models from nine providers. Every decision, allowed or denied, is written to a hash-chained evidence ledger. Agent request model, tools, budget Capability grant Ed25519-signed scope Router Thompson sampling 55 models 9 providers Evidence ledger hash-chained record of every allow and deny How BrainstormRouter handles an agent request: grant check, routing, models, and an evidence ledger for every decision Agent request model, tools, budget Capability grant Ed25519-signed scope Router Thompson sampling 55 models 9 providers Evidence ledger every allow and deny, hash-chained
Authority is checked before any model is called, and the ledger records denials as well as completed calls.
My commits
2,300+
Test files
1,400+
Stack
TypeScript
Origin
OpenClaw fork

brainstormrouter.com

Mar–Aug 2026

Brainstorm control plane

A governed control plane that connects AI coding assistants, such as Claude Code, to infrastructure through one channel. The question it answers: how do you give an assistant useful access to real systems while keeping its actions understandable and controllable?

  • MCP server with 58+ tools behind a single platform contract.
  • Approval workflows, cost tracking and audit reporting for agent actions.
  • Evaluation tooling for model output, a command-line interface and a desktop app.
My commits
2,000+
Packages
44
Test files
300+
Stack
TypeScript

brainstorm.co

Mar–Jun 2026

BrainstormLLM: learned orchestration

A small research project with a hard rule: ship only what beats the baseline. It predicts which phases of a software-development pipeline to run or skip for a given task, trained on 2,203 real Claude Code session turns, and exports to ONNX so the router can call it inline.

  • Sequential per-phase gradient-boosted models that condition each prediction on the phases before it.
  • Pre-set kill gates: mean F1 of at least 0.75, at least 25% cost reduction, under 10 ms inference. All three passed.
  • A separate model-tier classifier trained on 400,000+ RouterBench data points reached 93% cross-validated accuracy.
Mean F1
0.796
Cost reduction
68%
Inference
0.31 ms
Stack
Python, ONNX

brainstorm.co/llm

Apps

Native iPhone and web apps, built with the same agent workflow and taken through to real users.

Mar–Oct 2026
Built for Rebecca

Our Book Nook

A private reading journal, built for Rebecca, that a book club can share. Write a thought the moment it lands or say it out loud, bring in Kindle highlights with one tap, and keep it all private until you hand a book to a friend or open it with your club. Even then, nothing shows before they have read that far.

My commits
1,300+
Test files
430+
Stack
Swift, Next.js
Status
Live, free

ourbooknook.com and the App Store

Sep–Oct 2026
Built with Cooper

College Football Digest

A college football site Cooper and I are building together: scores, polls, team histories and headlines in one place, next to Cooper’s own takes. Cooper chose the design and writes the takes; I built the platform and the data behind it.

  • A DuckDB warehouse of every team, game, player and season, fed from CollegeFootballData.com behind a call-budget guard, that publishes page-ready data to the site.
  • Elo ratings built only from earlier games, which give each matchup a win probability on the home page.
  • Runs on Cloudflare Workers with D1 and R2, on the EmDash content system.
Commits
215 in 9 days
Cooper’s
29
Stack
Astro, Python
Data
DuckDB, D1

collegefootballdigest.com

Feb–Aug 2026

Peer10

A record of growing up through youth sports: a child's seasons, coach notes a parent can read and growth a kid can see, in one record that compounds year over year. League operations run simulation-first: intent, simulated change, human approval, execution and undo, the same agent-safety pattern as my other systems. Web and SwiftUI iPhone clients share one platform.

My commits
1,500+
Test files
550+
Stack
Next.js, SwiftUI
Data
Postgres, pgvector

peer10.com

Mar–Jul 2026

FinishStrong

Mobile-first SAT practice designed as a five-minute daily habit, built around an adaptive learning engine, on web and in a SwiftUI iPhone app.

My commits
220+
Stack
Next.js, SwiftUI
Data
Postgres

finishstrong.ai

Method

The common thread is a disciplined way of working with agents: they write most of the code, and I set the bar it has to clear.

  • Spec first

    Substantial work starts as a written specification with acceptance criteria, so an agent's output is judged against something other than its own description.

  • Gates, not vibes

    Results have to beat a baseline set in advance. BrainstormLLM shipped only after passing three pre-set kill gates.

  • Review loops

    Independent reviewer agents score the work against a stated bar, round after round, and a change lands only when the findings are fixed or explained.

  • Verify on the real path

    Done means proven where it runs: a live response, a database row, an App Store build, not a passing test in isolation.