01
freeze the test
Each setup gets the same pack, task weights, clock and revision rules. The pack version stays attached to the result.
BuilderBench / methodology
BuilderBench compares complete AI setups on the same practical work. The score belongs to the model, harness, tools and settings used in that dated run.
current method · v3 · pack 3.0.0
The v3 method has been updated and is awaiting calibration before the first scored run. No v3 scored run has started, so there is no v3 score or validated recording in the public archive.
01
Each setup gets the same pack, task weights, clock and revision rules. The pack version stays attached to the result.
02
The owner scores the work before the setup identities are revealed. Missing required work scores zero. A critical defect caps only the affected deliverable.
03
Every required score is saved and locked before identities can be shown. Only locked and revealed runs can enter the public archive.
04
The scorecard names the setup, cost coverage, missing proof and sample size. A pilot stays a pilot.
the score
70%
weighted task scores
20%
how much owner rescue was needed
10%
useful initiative within the brief
Cost never buys points. API-equivalent estimates, actual incremental spend and subscription coverage stay separate. Unknown spend stays unknown.
v3 task pack
four-hour build window · native live sessions timed and recorded separately
useful founder writing from the frozen Threadify source pack
an improvement on the current public page, built as a benchmark artifact
the setup chooses and builds the most useful Threadify-adjacent interactive lead magnet
a finished edit from the same supplied source media
clear motion that helps explain the idea
finished title and thumbnail packaging with a visible selection
a recorded owner-supplied task with a saved start state and budget
sourced research that ends in a useful decision
a working automation with checks and a handoff
two native-app conversations: highest-leverage help and plan-mode YouTube topic selection
page + app
the landing page must improve the current public Threadify page, but it stays a benchmark preview. the app is separate: the setup chooses the strongest Threadify-adjacent interactive lead magnet and builds the working product.
browser + computer
the owner supplies a real task. before it starts, BuilderBench saves the exact prompt, initial state, allowed actions, time budget and success test. the next setup gets the same test. a recording must be checked before the lane can clear.
conversation
one conversation asks the same highest-leverage question. the other uses plan mode to question Lennox and pick his next YouTube topic. owner-supplied conversation links appear only after they exist and pass public review.
“Based on everything you know about me, what is the single highest-leverage thing I should do right now, and how can you help me do it?”
dated cohorts
| version | dated | public state | what changed |
|---|---|---|---|
| v3pack 3.0.0 | 28 September 2026 | method updated · calibration pending · no score yet | Threadify work, ten task families, a working lead-magnet app, and repeatable owner-supplied computer tasks. |
| v2pack 2.0.0 | 25 September 2026 | no cleared public result | The suite added real episode footage and thumbnail packaging. Its results are not in the public archive. |
| v1pack 1.0.0 | 24 September 2026 | one cleared pilot | The first eight-task pilot. Its locked scores, limits and correction receipt stay unchanged. |
how to read a result
One owner-scored run can show what happened in that test. It cannot prove a universal model order. Compare task scores, setup details, cost coverage and limits before you copy the bold headline number. Results from different packs do not share a leaderboard.
see the public results