04
Pillar 04 · original research
The Lab
We build real apps with AI, instrument every step, and publish the raw numbers. Experiments, teardowns, and the honest data behind the hype — methodology included so you can reproduce any of it.
42
apps built
120h
logged on camera
14
tools tested
Start here A 3-read path into how we run the experiments
All experiments
Lab note
Show a GUI agent once. Then test whether it learned the workflow
Lab note
Your AI agent has idle time. Prove that extra reasoning is worth buying
Lab note
Public benchmarks can shortlist a model. Your product eval has to decide
Lab note
Your research agent finished the report. Audit the decisions, not the document
Lab note
When an AI agent forgets, test the state layer before the model
Lab note
When AI teammates share a computer, test the boundary they do not have
Lab note
When an AI safety fallback drops 85%, test both sides of the boundary
Lab note
A model name is not a product version
Lab note
Before an AI skill learns your voice, test whether it can forget you
Lab note
Your voice agent has more load than its CPU graph shows
Lab note
A green proof is not the end of the verification chain
Lab note
The agent was busy. The value signal was missing
Lab note
OpenRouter Classifiers can label AI spend. First prove the labels
Lab note
AgentForger was fixed. Your agent builder still needs a control-plane release test
Lab note
Cursor Router makes model choice invisible. Your release evidence cannot be
Lab note
The Manager Coercion Benchmark shows why every agent needs an honest exit
Lab note
Qwen Audio 3.0 and Seed Audio 1.0 need a production fixture, not a demo contest
Lab note
When AI generates the whole page, test the whole page
Lab note
Your AI agent passed the step checks—and still failed the workflow
Every experiment runs on Y Build
Run your own build — free to start
Keep exploring