In the Stage 2 report we shipped 20 feature tickets on Fizzy and promised to explore benchmarking the agents all on max-effort. Now we have run it: every model on the board, same tickets, reasoning turned all the way up. We also tested a new model: DeepSeek 4.1 Flash.
The short version:
- Effort matters, but only with some agents. GPT-6 Astra (already in first place) went from 35% to 53% success, GPT-5.6 Sol improved from 18% to 28%, and GPT-5.6 Luna performed much better: from 0% to 27% success. Claude, Gemini and Grok barely moved, and Gemini regressed. Muse doubled, from 10% to 20%.
- It costs. The max sweep cost about $4,100 across all models against $2,250 for the defaults, and runs…









An Abyssinian Blue sits at a laptop, paw on the trackpad planning a dire conspiracy.









