A reproducible comparison
Every model receives the same real-world task and constraints.
Temperature zero and a sandboxed workflow keep every attempt comparable.
Review quality, cost and latency together—not a vendor-picked benchmark.
Benchmarks are self-reported — vendors pick the tests, tune the prompts, and publish the numbers. The ASS Whole Outline is the opposite: fixed prompts, temperature zero, identical scoring, run in a sandbox with no cherry-picking and no re-rolls. Every number here is a live run we can reproduce, not a claim we have to trust.