Lennox Saint
Open menu
pack 3.0.0

30 September 2026

1 owner score per setup

owner-finalized partial personal assessment

BuilderBench · v3

GPT-6.1 Sol scored 41.75.

One owner-finalized partial review of one GPT-6.1 Sol run with Codex on the v3 pack. This is not a universal model ranking or a statistical claim.

owner-finalized partial personal assessment. not a universal model ranking.

BuilderBench v3 result: GPT-6.1 Sol with Codex scored 41.75 out of 100 (quality 41.07, autonomy 30, proactivity 70) in an owner-finalized partial review. API-equivalent estimate US$14.71, partial usage.
owner-finalized partial personal assessment. costs are estimates, not bills. cost coverage is shown beside each figure.

overall score

quality 70% · autonomy 20% · proactivity 10%

Codex harness · pack 3.0.0

GPT-6.1 Sol

41.75/100

quality 70% of the score
41.07
autonomy 20%
30
proactivity 10%
70

task scores

where it gained and lost ground

each task is scored out of 100. its weight is its share of the quality score.

  1. Writing · weight 10%

    GPT-6.1 Sol26
  2. Landing page · weight 10%

    GPT-6.1 Sol30
  3. App lead magnet · weight 20%

    GPT-6.1 Sol30
  4. Video editing · weight 10%

    GPT-6.1 Sol16.7
  5. Motion · weight 5%

    GPT-6.1 Sol20
  6. Thumbnails · weight 10%

    GPT-6.1 Sol23
  7. Browser + computer · weight 10%

    GPT-6.1 Sol55
  8. Research · weight 10%

    GPT-6.1 Sol55
  9. Automation · weight 10%

    GPT-6.1 Sol100
  10. Founder conversations · weight 5%

    GPT-6.1 Sol70
see the task table
Task scores, weights and API-equivalent cost estimates for each setup
taskweightGPT-6.1 Sol Codex · scorecost API-equivalent estimate
Writing10%26US$0.50recorded
Landing page10%30US$1.07partly recorded
App lead magnet20%30US$4.13recorded
Video editing10%16.67US$5.32recorded
Motion5%20US$1.31partly recorded
Thumbnails10%23US$1.71recorded
Browser + computer10%55unknownnot recorded
Research10%55US$0.22partly recorded
Automation10%100US$0.45recorded
Founder conversations5%70unknownnot recorded

frozen setup

what was tested

pack version
3.0.0
clock
4 hours
reasoning
Extra High (xhigh)
speed
Standard
billing
Existing subscriptions
owner scores
1 per setup

exact model and harness version strings are not in this public export, so they stay unknown here.

cost

estimates are not bills

GPT-6.1 Sol

US$14.71

API-equivalent estimate for recorded execution usage; 15 of 18 execution records priced.

15 of 18 execution records priced

actual incremental spend: unknown

API-equivalent estimates price recorded execution usage at published API rates. They are not bills. Missing native usage, setup, calibration, coordination, grading, presentation, subscription fees and unreceipted services are excluded. Actual incremental spend is unknown.

limits

what this result does not prove

  • This is one owner-finalized partial personal assessment of one run, not a universal ranking.
  • Independent checks and native recordings did not finish. Saved outputs were assessed as delivered; the original execution and evidence remain incomplete.
  • The founder conversations ran in the owner's own Codex app with app memory on, not from the fixture-only source pack. The planning conversation ran in default mode, not plan mode, and its rating was saved before that conversation ended.
  • The score belongs to the full setup: model, harness, tools and settings.
  • Pack 3.0.0 scores compare only with other pack 3.0.0 results, not with v1 or v2 results.
  • The API-equivalent estimate covers 15 of 18 execution records. Tasks marked not recorded have no priced usage.
  • Actual incremental spend is unknown because the run used existing subscriptions.
  • The build window was four hours. Native live sessions were timed separately from the build window.

The owner rated the saved outputs of each of the ten v3 tasks, plus autonomy and proactivity, in a named single-model review. The saved scores were then finalized as a partial personal assessment and locked. Quality is the weighted task score. v1 to v3 use 70% quality, 20% autonomy and 10% proactivity.

p.s.

get the next BuilderBench result first.

  • the exact prompt i typed
  • what came back
  • where it fell over

free. unsubscribe any time, one click. privacy.

  • 32,605 on Threads
  • 3,750 readers
  • 1.79K on YouTube

live count, 6 oct 2026 / 7 oct 2026.