Lennox Saint
Open menu

BuilderBench / methodology

judge the work. keep the receipt.

BuilderBench compares complete AI setups on the same practical work. The score belongs to the model, harness, tools and settings used in that dated run.

current method · v3 · pack 3.0.0

no score yet

The v3 method has been updated and is awaiting calibration before the first scored run. No v3 scored run has started, so there is no v3 score or validated recording in the public archive.

01

freeze the test

Each setup gets the same pack, task weights, clock and revision rules. The pack version stays attached to the result.

02

review blind

The owner scores the work before the setup identities are revealed. Missing required work scores zero. A critical defect caps only the affected deliverable.

03

lock, then reveal

Every required score is saved and locked before identities can be shown. Only locked and revealed runs can enter the public archive.

04

publish the limits

The scorecard names the setup, cost coverage, missing proof and sample size. A pilot stays a pilot.

the score

quality carries the most weight

70%

quality

weighted task scores

20%

autonomy

how much owner rescue was needed

10%

proactivity

useful initiative within the brief

Cost never buys points. API-equivalent estimates, actual incremental spend and subscription coverage stay separate. Unknown spend stays unknown.

v3 task pack

ten task families. 100 points.

four-hour build window · native live sessions timed and recorded separately

  1. 01

    writing

    10

    useful founder writing from the frozen Threadify source pack

  2. 02

    improved Threadify landing page

    10

    an improvement on the current public page, built as a benchmark artifact

  3. 03

    app lead magnet

    20

    the setup chooses and builds the most useful Threadify-adjacent interactive lead magnet

  4. 04

    video

    10

    a finished edit from the same supplied source media

  5. 05

    motion

    5

    clear motion that helps explain the idea

  6. 06

    thumbnails

    10

    finished title and thumbnail packaging with a visible selection

  7. 07

    browser + computer

    10

    a recorded owner-supplied task with a saved start state and budget

  8. 08

    research

    10

    sourced research that ends in a useful decision

  9. 09

    automation

    10

    a working automation with checks and a handoff

  10. 10

    conversation

    5

    two native-app conversations: highest-leverage help and plan-mode YouTube topic selection

page + app

useful work, not a costume

the landing page must improve the current public Threadify page, but it stays a benchmark preview. the app is separate: the setup chooses the strongest Threadify-adjacent interactive lead magnet and builds the working product.

browser + computer

save the start, then record the work

the owner supplies a real task. before it starts, BuilderBench saves the exact prompt, initial state, allowed actions, time budget and success test. the next setup gets the same test. a recording must be checked before the lane can clear.

conversation

test the native app twice

one conversation asks the same highest-leverage question. the other uses plan mode to question Lennox and pick his next YouTube topic. owner-supplied conversation links appear only after they exist and pass public review.

“Based on everything you know about me, what is the single highest-leverage thing I should do right now, and how can you help me do it?”

dated cohorts

new work or weights means a new version

versiondatedpublic statewhat changed
v3pack 3.0.028 September 2026method updated · calibration pending · no score yetThreadify work, ten task families, a working lead-magnet app, and repeatable owner-supplied computer tasks.
v2pack 2.0.025 September 2026no cleared public resultThe suite added real episode footage and thumbnail packaging. Its results are not in the public archive.
v1pack 1.0.024 September 2026one cleared pilotThe first eight-task pilot. Its locked scores, limits and correction receipt stay unchanged.

how to read a result

a scorecard is a dated experiment

One owner-scored run can show what happened in that test. It cannot prove a universal model order. Compare task scores, setup details, cost coverage and limits before you copy the bold headline number. Results from different packs do not share a leaderboard.

see the public results