i give AI my real work.
same jobs. same clock. i score the work blind. then i publish the result, the limits and the corrections.
latest cleared result · 24 September 2026 · pack 1.0.0
Opus 5.5 won. by a lot.
quality 70% · autonomy 20% · proactivity 10%. out of 100.

Opus 5.5with Claude
68.26/100

GPT-6 Solwith Codex
45.55/100
pack 3.0.0
no score yet
the next cohort
ten real jobs built around Threadify.
v3 uses useful Threadify source material. the landing-page task improves the current public page as a benchmark preview. it does not replace production. the new app task makes each setup choose and build a working Threadify-adjacent lead magnet.
the build window is four hours. native live sessions are timed and recorded separately from the build window. the v3 method has been updated and is awaiting calibration before the first scored run. no v3 scored run has started, so there is no v3 score or validated recording in the public archive.
see all ten taskshow it works
four rules. no exceptions.

real jobs, not puzzles.
a frozen pack of work i actually do: writing, landing pages, browser jobs, workflow fixes.

BuilderBench v1 
every job is weighted.
each job counts for a set share. the scores roll up into one number out of 100.

from LAB 0009 
scored blind. then locked.
i score the work before i know which AI did it. then the scores lock. no take-backs.

from LAB 0009 
costs and mistakes go public too.
every score, what it cost to run, the limits and any correction. dated.

from LAB 0009

the point: your work is the benchmark.
steal the method. score the AI on your own jobs.
read the full methodwhat this doesn’t prove
- This was one pilot scored by one owner, not a universal ranking.
- The score belongs to each full setup: model, harness, tools and settings.
- The two cost estimates use different methods and coverage.
- Actual incremental spend remains unknown because the run used existing subscriptions.
- Complete normal-speed watched-and-listened review of both full video outputs was not established.
how the scoring works
each setup gets the same frozen pack and the same clock. i rate the finished work before identities are revealed, then lock the scores. one pilot is not a universal ranking.
read the full methodversion history
a changed job or weight starts a new cohort.
v3 · pack 3.0.0
28 September 2026
method updated · calibration pending · no score yet
Threadify work, ten task families, a working lead-magnet app, and repeatable owner-supplied computer tasks.
v2 · pack 2.0.0
25 September 2026
no cleared public result
The suite added real episode footage and thumbnail packaging. Its results are not in the public archive.
v1 · pack 1.0.0
24 September 2026
one cleared pilot
The first eight-task pilot. Its locked scores, limits and correction receipt stay unchanged.
open the cleared result
changelog
every change, dated. corrections get their own line.
method· v3
v3 method: ten task families on real Threadify work
Threadify work, ten task families, a working lead-magnet app, and repeatable owner-supplied computer tasks. Calibration is pending, so there is no v3 score yet.
open itnew cohort· v2
v2: real episode footage and thumbnail packaging
The suite added real episode footage and thumbnail packaging. Its results are not in the public archive.
new result· v1
v1 pilot result: Opus 5.5 68.26, GPT-6 Sol 45.55
Opus 5.5 with Claude scored 68.26 and GPT-6 Sol with Codex scored 45.55. One owner-scored run per setup, so it is not a universal model ranking.
open itcorrection· v1
correction before release: stale missing-file flags
After the scores were locked, live-chat deliverables were still flagged as missing. The frozen files and their hashes were checked again, and they were there. My ratings of the work did not change. Only the missing-file flags did.
open itnew cohort· v1
v1: the first eight-task pilot
The first eight-task pilot: two AI setups, the same frozen jobs, the same four-hour clock, scored blind.
open it
cleared results
grouped by task pack. different packs don’t compare.
p.s.
get the next BuilderBench result first.
- the exact prompt i typed
- what came back
- where it fell over
- 32,545 on Threads
- 3,752 readers
- 1.78K on YouTube
live count, 29 sep 2026.













