Building
What I am making, what I am learning in order to make it, and the tools I make it with.
Projects
Two public repositories you can run, and one system I keep private.
A replay engine for AI agent runs. It re-derives the steps a machine can re-derive and reports match, divergence, or could not check. The third outcome is the point: a validator that cannot tell a lie from a blind spot is not a validator.
A tutor that reads my notes, asks real questions, adapts at runtime, and settles a signed proof of what I studied on Ethereum Sepolia. Compute off chain, verify on chain.
Five agent systems that run my days, each shaped around me rather than an average user. Built on the Claude Agent SDK, with scheduled tasks and sub-agents.
Latest commits
Pulled from GitHub every few hours, so this list moves when the work moves.
- 2026-10-04proof-of-agent-rundocs: update README, worklog for session five
- 2026-10-04proof-of-agent-runfeat: make the engine installable
- 2026-10-04proof-of-agent-rundocs: shorten docstrings, keep design reasoning in TRACE-FORMAT
- 2026-10-04proof-of-agent-runfeat: guard table is a parameter and each entry names its answer key
- 2026-09-28proof-of-agent-rundocs: worklog for the wrong-type-containment session
- 2026-09-28proof-of-agent-rundocs: catch clause 3 and the BAD_TYPE_RUN comments up to the wrong_type containment
What I am learning right now
The ground under the build. One concept at a time, and each one ends up in an episode.
Consensus, gossip and the mempool, from the ground up. What is public before anything is ordered.
State, the trie, and the root that makes agreement between thousands of machines cheap.
Four ways to check a stranger's AI work: cryptographic proofs, trusted hardware, reproducible execution, staked validators.
ERC-8004 registries, and the empty slot under agent validation that my replay engine is built for.
Distributed training, and the communication wall between machines that keeps it expensive.
The trace format for proof-of-agent-run: what a machine can honestly re-derive about another machine's run.
Stack
What I reach for when an agent has to work for real.
My default for everything that has to run in production: integrations, agents, automation and the build scripts around them.
Skills, scheduled tasks and sub-agent orchestration. The layer most of my agent systems stand on.
For agents that have to pause for hours between turns and pick up exactly where they stopped.
Contracts and their tests, on testnet. Where the oracle pattern gets written down as code.
The telephony underneath voice agents: numbers, routing, voicemail detection and the quirks of real calls.
How my Python agents reach the chain: signing attestations, submitting them, and reading the result back.
Get in touch
Work, an idea, or a correction to something I said on the show. This is where to find me, and I reply.