Built on Y Build Every experiment runs on Y Build — go from prompt to a deployed app, no server. Start free
BuildShipCompareThe LabAbout Start building →
04
Pillar 04 · original research

The Lab

We build real apps with AI, instrument every step, and publish the raw numbers. Experiments, teardowns, and the honest data behind the hype — methodology included so you can reproduce any of it.

42
apps built
120h
logged on camera
14
tools tested
Start here A 3-read path into how we run the experiments

All experiments

Lab note
Show a GUI agent once. Then test whether it learned the workflow
Aug 19, 2026 · Jordan Park · 19 min
Read →
Lab note
Your AI agent has idle time. Prove that extra reasoning is worth buying
Aug 18, 2026 · Maya Chen · 18 min
Read →
Lab note
Public benchmarks can shortlist a model. Your product eval has to decide
Aug 17, 2026 · Alex Liu · 19 min
Read →
Lab note
Your research agent finished the report. Audit the decisions, not the document
Aug 15, 2026 · Noah Bennett · 19 min
Read →
Lab note
When an AI agent forgets, test the state layer before the model
Aug 13, 2026 · Elena Torres · 18 min
Read →
Lab note
When AI teammates share a computer, test the boundary they do not have
Aug 12, 2026 · Jordan Park · 19 min
Read →
Lab note
When an AI safety fallback drops 85%, test both sides of the boundary
Aug 8, 2026 · Maya Chen · 19 min
Read →
Lab note
A model name is not a product version
Aug 7, 2026 · Alex Liu · 18 min
Read →
Lab note
Before an AI skill learns your voice, test whether it can forget you
Aug 6, 2026 · Noah Bennett · 19 min
Read →
Lab note
Your voice agent has more load than its CPU graph shows
Aug 4, 2026 · Elena Torres · 18 min
Read →
Lab note
A green proof is not the end of the verification chain
Aug 2, 2026 · Jordan Park · 18 min
Read →
Lab note
The agent was busy. The value signal was missing
Jul 31, 2026 · Maya Chen · 18 min
Read →
Lab note
OpenRouter Classifiers can label AI spend. First prove the labels
Jul 25, 2026 · Alex Liu · 18 min
Read →
Lab note
AgentForger was fixed. Your agent builder still needs a control-plane release test
Jul 24, 2026 · Noah Bennett · 19 min
Read →
Lab note
Cursor Router makes model choice invisible. Your release evidence cannot be
Jul 23, 2026 · Elena Torres · 19 min
Read →
Lab note
The Manager Coercion Benchmark shows why every agent needs an honest exit
Jul 22, 2026 · Jordan Park · 18 min
Read →
Lab note
Qwen Audio 3.0 and Seed Audio 1.0 need a production fixture, not a demo contest
Jul 21, 2026 · Maya Chen · 18 min
Read →
Lab note
When AI generates the whole page, test the whole page
Jul 20, 2026 · Alex Liu · 18 min
Read →
Lab note
Your AI agent passed the step checks—and still failed the workflow
Jul 19, 2026 · Y Build Editorial · 19 min
Read →
Every experiment runs on Y Build
Run your own build — free to start
Start building
Keep exploring
Run your own build
Free · no card
Start free →