Lennox Saint
Open menu
pack 3.0.0

29 September 2026

1 owner score per setup

owner-finalized partial personal assessment

BuilderBench · v3

Sonnet 5.5 scored 48.98.

One owner-finalized partial review of one Sonnet 5.5 run with Claude on the v3 pack. This is not a universal model ranking or a statistical claim.

owner-finalized partial personal assessment. not a universal model ranking.

BuilderBench v3 result: Sonnet 5.5 with Claude scored 48.98 out of 100 (quality 51.4, autonomy 40, proactivity 50) in an owner-finalized partial review. API-equivalent estimate US$17.59, partial usage.
owner-finalized partial personal assessment. costs are estimates, not bills. cost coverage is shown beside each figure.

overall score

quality 70% · autonomy 20% · proactivity 10%

Claude harness · pack 3.0.0

Sonnet 5.5

48.98/100

quality 70% of the score
51.40
autonomy 20%
40
proactivity 10%
50

task scores

where it gained and lost ground

each task is scored out of 100. its weight is its share of the quality score.

  1. Writing · weight 10%

    Sonnet 5.536
  2. Landing page · weight 10%

    Sonnet 5.545
  3. App lead magnet · weight 20%

    Sonnet 5.542.5
  4. Video editing · weight 10%

    Sonnet 5.550
  5. Motion · weight 5%

    Sonnet 5.555
  6. Thumbnails · weight 10%

    Sonnet 5.560
  7. Browser + computer · weight 10%

    Sonnet 5.548
  8. Research · weight 10%

    Sonnet 5.550
  9. Automation · weight 10%

    Sonnet 5.585
  10. Founder conversations · weight 5%

    Sonnet 5.555
see the task table
Task scores, weights and API-equivalent cost estimates for each setup
taskweightSonnet 5.5 Claude · scorecost API-equivalent estimate
Writing10%36US$1.49recorded
Landing page10%45US$3.58partly recorded
App lead magnet20%42.50unknownnot recorded
Video editing10%50US$4.09partly recorded
Motion5%55US$3.26partly recorded
Thumbnails10%60US$2.90recorded
Browser + computer10%48unknownnot recorded
Research10%50US$0.58recorded
Automation10%85US$1.69recorded
Founder conversations5%55unknownnot recorded
correction

the correction is a receipt

2026-10-07 · fixed after this result went public

sonnet 5.5 v3 graphic date: The Sonnet 5.5 v3 result graphic shows 30 September 2026. That is the day the partial review was finalized. The run itself was on 29 September 2026, as the result page says. Scores, costs and limits are unchanged.

frozen setup

what was tested

pack version
3.0.0
clock
4 hours
reasoning
Extra High (xhigh)
speed
Standard
billing
Existing subscriptions
owner scores
1 per setup

exact model and harness version strings are not in this public export, so they stay unknown here.

cost

estimates are not bills

Sonnet 5.5

US$17.59

API-equivalent estimate for recorded execution usage; 9 of 14 execution records priced.

9 of 14 execution records priced

actual incremental spend: unknown

API-equivalent estimates price recorded execution usage at published API rates. They are not bills. Missing native usage, setup, calibration, coordination, grading, presentation, subscription fees and unreceipted services are excluded. Actual incremental spend is unknown.

limits

what this result does not prove

  • This is one owner-finalized partial personal assessment of one run, not a universal ranking.
  • Execution was interrupted. Saved outputs were assessed as delivered; the original execution and evidence remain incomplete.
  • The score belongs to the full setup: model, harness, tools and settings.
  • Pack 3.0.0 scores compare only with other pack 3.0.0 results, not with v1 or v2 results.
  • The API-equivalent estimate covers 9 of 14 execution records. Tasks marked not recorded have no priced usage.
  • Actual incremental spend is unknown because the run used existing subscriptions.
  • The build window was four hours. Native live sessions were timed separately from the build window.

The owner rated the saved outputs of each of the ten v3 tasks, plus autonomy and proactivity, in a named single-model review. The saved scores were then finalized as a partial personal assessment and locked. Quality is the weighted task score. v1 to v3 use 70% quality, 20% autonomy and 10% proactivity.

p.s.

get the next BuilderBench result first.

  • the exact prompt i typed
  • what came back
  • where it fell over

free. unsubscribe any time, one click. privacy.

  • 32,605 on Threads
  • 3,750 readers
  • 1.79K on YouTube

live count, 6 oct 2026 / 7 oct 2026.