Lennox Saint
Open menu
pack 2.0.0

26 September 2026

1 owner score per setup

personal assessment with accepted recording limitations

BuilderBench · v2

GPT-6 Astra scored 53.75.

One owner-scored personal assessment of one GPT-6 Astra run with Codex on the v2 pack. This is not a universal model ranking or a statistical claim.

personal assessment with accepted recording limitations. not a universal model ranking.

BuilderBench v2 result: GPT-6 Astra with Codex scored 53.75 out of 100 (quality 52.5, autonomy 60, proactivity 50) in a personal assessment with accepted recording limitations. API-equivalent estimate US$87.50, partial usage.
personal assessment with accepted recording limitations. costs are estimates, not bills. cost coverage is shown beside each figure.

overall score

quality 70% · autonomy 20% · proactivity 10%

Codex harness · pack 2.0.0

GPT-6 Astra

53.75/100

quality 70% of the score
52.50
autonomy 20%
60
proactivity 10%
50

task scores

where it gained and lost ground

each task is scored out of 100. its weight is its share of the quality score.

  1. Writing · weight 15%

    GPT-6 Astra32
  2. Video editing · weight 15%

    GPT-6 Astra43.3
  3. Landing page · weight 15%

    GPT-6 Astra35
  4. Motion · weight 5%

    GPT-6 Astra20
  5. Browser + computer · weight 15%

    GPT-6 Astra90
  6. Research · weight 10%

    GPT-6 Astra50
  7. Automation · weight 10%

    GPT-6 Astra80
  8. Founder conversations · weight 5%

    GPT-6 Astra65
  9. Thumbnails · weight 10%

    GPT-6 Astra52
see the task table
Task scores, weights and API-equivalent cost estimates for each setup
taskweightGPT-6 Astra Codex · scorecost API-equivalent estimate
Writing15%32US$2.90recorded
Video editing15%43.33US$28.81partly recorded
Landing page15%35US$15.25recorded
Motion5%20US$16.32recorded
Browser + computer15%90US$7.55recorded
Research10%50US$3.16recorded
Automation10%80US$2.47recorded
Founder conversations5%65US$0.89recorded
Thumbnails10%52US$10.15partly recorded

frozen setup

what was tested

pack version
2.0.0
clock
4 hours
reasoning
Extra High (xhigh)
speed
Standard
billing
Existing subscriptions
owner scores
1 per setup

exact model and harness version strings are not in this public export, so they stay unknown here.

cost

estimates are not bills

GPT-6 Astra

US$87.50

API-equivalent estimate for recorded execution usage; 20 of 31 execution records priced.

20 of 31 execution records priced

actual incremental spend: unknown

API-equivalent estimates price recorded execution usage at published API rates. They are not bills. Missing native usage, setup, calibration, coordination, grading, presentation, subscription fees and unreceipted services are excluded. Actual incremental spend is unknown.

limits

what this result does not prove

  • This is one owner-scored personal assessment of one run, not a universal ranking.
  • The original browser and computer recordings have incomplete action coverage. The owner accepted this limit for personal scoring using the available recordings and saved outputs.
  • The score belongs to the full setup: model, harness, tools and settings.
  • Pack 2.0.0 scores compare only with other pack 2.0.0 results, not with v1 or v3 results.
  • The API-equivalent estimate covers 20 of 31 execution records. Tasks marked not recorded have no priced usage.
  • Actual incremental spend is unknown because the run used existing subscriptions.
  • The build window was four hours.

The owner rated the saved outputs of each of the nine v2 tasks, plus autonomy and proactivity, then locked the scores. An accepted evidence limitation covers the browser and computer task. Quality is the weighted task score. v1 to v3 use 70% quality, 20% autonomy and 10% proactivity.

p.s.

get the next BuilderBench result first.

  • the exact prompt i typed
  • what came back
  • where it fell over

free. unsubscribe any time, one click. privacy.

  • 32,605 on Threads
  • 3,750 readers
  • 1.79K on YouTube

live count, 6 oct 2026 / 7 oct 2026.