Claude harness · pack 3.0.0
Sonnet 5.5
48.98/100
- quality 70% of the score
- 51.40
- autonomy 20%
- 40
- proactivity 10%
- 50
29 September 2026
1 owner score per setup
owner-finalized partial personal assessment

BuilderBench · v3
One owner-finalized partial review of one Sonnet 5.5 run with Claude on the v3 pack. This is not a universal model ranking or a statistical claim.
owner-finalized partial personal assessment. not a universal model ranking.

overall score
Claude harness · pack 3.0.0
48.98/100
task scores
each task is scored out of 100. its weight is its share of the quality score.
Writing · weight 10%
Landing page · weight 10%
App lead magnet · weight 20%
Video editing · weight 10%
Motion · weight 5%
Thumbnails · weight 10%
Browser + computer · weight 10%
Research · weight 10%
Automation · weight 10%
Founder conversations · weight 5%
| task | weight | Sonnet 5.5 Claude · score | cost API-equivalent estimate |
|---|---|---|---|
| Writing | 10% | 36 | US$1.49recorded |
| Landing page | 10% | 45 | US$3.58partly recorded |
| App lead magnet | 20% | 42.50 | unknownnot recorded |
| Video editing | 10% | 50 | US$4.09partly recorded |
| Motion | 5% | 55 | US$3.26partly recorded |
| Thumbnails | 10% | 60 | US$2.90recorded |
| Browser + computer | 10% | 48 | unknownnot recorded |
| Research | 10% | 50 | US$0.58recorded |
| Automation | 10% | 85 | US$1.69recorded |
| Founder conversations | 5% | 55 | unknownnot recorded |
2026-10-07 · fixed after this result went public
sonnet 5.5 v3 graphic date: The Sonnet 5.5 v3 result graphic shows 30 September 2026. That is the day the partial review was finalized. The run itself was on 29 September 2026, as the result page says. Scores, costs and limits are unchanged.
frozen setup
exact model and harness version strings are not in this public export, so they stay unknown here.
cost
US$17.59
API-equivalent estimate for recorded execution usage; 9 of 14 execution records priced.
9 of 14 execution records priced
actual incremental spend: unknown
API-equivalent estimates price recorded execution usage at published API rates. They are not bills. Missing native usage, setup, calibration, coordination, grading, presentation, subscription fees and unreceipted services are excluded. Actual incremental spend is unknown.
limits
The owner rated the saved outputs of each of the ten v3 tasks, plus autonomy and proactivity, in a named single-model review. The saved scores were then finalized as a partial personal assessment and locked. Quality is the weighted task score. v1 to v3 use 70% quality, 20% autonomy and 10% proactivity.
p.s.
live count, 6 oct 2026 / 7 oct 2026.