Lennox Saint
Open menu
pack 1.0.0

24 September 2026

1 owner score per setup

BuilderBench · v1 pilot

Opus 5.5 scored 68.26. GPT-6 Sol scored 45.55.

One owner-scored pilot of each model, harness and tool setup. This is not a universal model ranking or a statistical claim.

one owner-scored pilot. not a universal model ranking.

BuilderBench pilot result: Opus 5.5 scored 68.26 and GPT-6 Sol scored 45.55, with API-equivalent estimate caveats.
one owner-scored pilot. costs are estimates, not bills. Sol usage is partly recorded. methods and coverage differ.

overall score

quality 70% · autonomy 20% · proactivity 10%

Claude harness

Opus 5.5

68.26/100

quality
61.8
autonomy
80
proactivity
90

Codex harness

GPT-6 Sol

45.55/100

quality
46.5
autonomy
40
proactivity
50

task scores

where each setup gained and lost ground

  1. Writing · weight 20%

    Opus 5.549
    GPT-6 Sol30
  2. Video editing · weight 15%

    Opus 5.540
    GPT-6 Sol10
  3. Frontend design · weight 15%

    Opus 5.560
    GPT-6 Sol40
  4. Motion graphics · weight 10%

    Opus 5.560
    GPT-6 Sol30
  5. Browser & computer · weight 15%

    Opus 5.570
    GPT-6 Sol90
  6. Research & decisions · weight 10%

    Opus 5.570
    GPT-6 Sol55
  7. Intake automation · weight 10%

    Opus 5.595
    GPT-6 Sol90
  8. Founder conversation · weight 5%

    Opus 5.580
    GPT-6 Sol40
see the task table
Task scores and weights for each setup
taskweightOpus 5.5 ClaudeGPT-6 Sol Codex
Writing20%4930
Video editing15%4010
Frontend design15%6040
Motion graphics10%6030
Browser & computer15%7090
Research & decisions10%7055
Intake automation10%9590
Founder conversation5%8040
correction

the correction is a receipt

2026-09-24 15:13 UTC · fixed before this result went public

what happened: after the scores were locked, live-chat deliverables were still flagged as missing.

why: stale missing-file flags from the offline window. the frozen files and their hashes were checked again, and they were there.

what did not change: my ratings of the work did not change. only the missing-file flags did.

Scores before and after the correction
setupscorebeforeafter
Opus 5.5 · Claudequality57.8061.80
Opus 5.5 · Claudeoverall65.4668.26
GPT-6 Sol · Codexquality44.5046.50
GPT-6 Sol · Codexoverall44.1545.55

frozen setup

what was tested

pack version
1.0.0
clock
4 hours
reasoning
Extra High (xhigh)
speed
Standard
billing
Existing subscriptions
owner scores
1 per setup

exact model and harness version strings were not kept in the public pilot export, so they stay unknown here.

cost

estimates are not bills

Opus 5.5

US$49.67

Provider-reported API estimates for the recorded task rows.

GPT-6 Sol

US$12.75

Partial recorded usage only; unknown attempts and shared overhead are excluded.

API-equivalent estimates are not bills. Actual incremental spend is unknown. The methods and coverage differ, so these figures do not support an exact cost ratio.

limits

what this pilot does not prove

  • This was one pilot scored by one owner, not a universal ranking.
  • The score belongs to each full setup: model, harness, tools and settings.
  • The two cost estimates use different methods and coverage.
  • Actual incremental spend remains unknown because the run used existing subscriptions.
  • Complete normal-speed watched-and-listened review of both full video outputs was not established.

The owner scored quality, autonomy and proactivity after a blind review. Quality is the weighted score across eight task families. The score lock happened before identities were revealed.

p.s.

get the next BuilderBench result first.

  • the exact prompt i typed
  • what came back
  • where it fell over

free. unsubscribe any time, one click. privacy.

  • 32,545 on Threads
  • 3,752 readers
  • 1.78K on YouTube

live count, 29 sep 2026.