The archive
All field notes
Every build, ship guide, comparison, and experiment — newest first.
★ FLAGSHIP
04 · The Lab
Show a GUI agent once. Then test whether it learned the workflow
★ FLAGSHIP
04 · The Lab
Your AI agent has idle time. Prove that extra reasoning is worth buying
★ FLAGSHIP
04 · The Lab
Public benchmarks can shortlist a model. Your product eval has to decide
★ FLAGSHIP
04 · The Lab
Your research agent finished the report. Audit the decisions, not the document
★ FLAGSHIP
04 · The Lab
When an AI agent forgets, test the state layer before the model
★ FLAGSHIP
04 · The Lab
When AI teammates share a computer, test the boundary they do not have
★ FLAGSHIP
04 · The Lab
When an AI safety fallback drops 85%, test both sides of the boundary
★ FLAGSHIP
04 · The Lab
A model name is not a product version
★ FLAGSHIP
04 · The Lab
Before an AI skill learns your voice, test whether it can forget you
★ FLAGSHIP
04 · The Lab
Your voice agent has more load than its CPU graph shows
★ FLAGSHIP
04 · The Lab
A green proof is not the end of the verification chain
★ FLAGSHIP
04 · The Lab
The agent was busy. The value signal was missing
★ FLAGSHIP
04 · The Lab
OpenRouter Classifiers can label AI spend. First prove the labels
★ FLAGSHIP
04 · The Lab
AgentForger was fixed. Your agent builder still needs a control-plane release test
★ FLAGSHIP
04 · The Lab
Cursor Router makes model choice invisible. Your release evidence cannot be
★ FLAGSHIP
04 · The Lab
The Manager Coercion Benchmark shows why every agent needs an honest exit
★ FLAGSHIP
04 · The Lab
Qwen Audio 3.0 and Seed Audio 1.0 need a production fixture, not a demo contest
★ FLAGSHIP
04 · The Lab
When AI generates the whole page, test the whole page
★ FLAGSHIP
04 · The Lab
Your AI agent passed the step checks—and still failed the workflow
02 · Ship
Deploy an AI-built app to your own domain
02 · Ship
Custom domains & HTTPS without touching DNS
Reading is research. Building is faster.
Start your first app free — deploy to your own domain in one pass.