Every model, one reference.
Where most teams start
Anthropic's new flagship: 80.3% SWE-bench Pro, 96% SWE-bench Verified on Vals.ai, and 85.0% OSWorld-Verified make it the best production coding pick for non-trivial engineering tasks.
Everyday agent at Sonnet list $3 / $15 per 1M tokens, with a 1M window, 81.2% OSWorld-Verified, and 86.6% BrowseComp multi-agent.
Tops Chatbot Arena (1503) and writes paragraphs you'd ship; understands tone notes and edits like a copy chief.
GDPval-AA ELO 1932 and Anthropic-reported finance, trading, and analytics wins make it the strongest general knowledge-work pick; do not use Mythos-only HLE rows as Fable evidence.
FLUX.2 Klein 9B is a 9B distilled text-to-image model on fal. Not FLUX.2 Dev.
Best overall video quality in the catalog: 30-second clips, native audio, and up to 4K through Vertex AI.