live
since sep 2026
public lab
Lenny’s AI Builders + BuilderBench
the AI catch-up plus one thing i actually tried - with public field notes and a benchmark that shows its working.
open it
what i did
- launched LAB 0000 on 17 sep 2026. seven episodes in the first eight days.
- wrote public field notes for each one: the steps, the prompt, the limits.
- built BuilderBench: two AI setups, the same work, the same four-hour clock, scored by me.
what changed
- the v1 pilot scored Opus 5.5 with Claude 68.26 and GPT-6 Sol with Codex 45.55.
- after the scores were locked, i found a scoring mistake, fixed it, and dated the correction on the scorecard.
what it proves
i test the tools on real work and show the receipts, including the mistakes.
what it doesn’t
which model is best for you. one owner-scored pilot is not a ranking.