In September 2026, Anthropic's AI models claimed seven of the top ten spots in the Arena rankings, a user-driven benchmark that evaluates the performance of artificial intelligence systems. The company, known for its Claude series, placed its latest model, Claude Fable 5, at the top of the list. This was the first public release of the Mythos class of models, which Anthropic described as a significant advancement in AI capabilities. Two earlier versions, Claude Opus 4.6 and Claude Opus 4.7, also ranked highly, showing that older models still held strong in the competition. However, the newest Claude Fable 5.1, launched at the start of September, did not surpass fifth place. It was overtaken by Muse Spark 1.2, a model developed by Meta, which powers the Muse Code agent. Meta also placed Muse Spark 1.3 in eighth place, edging out Google's Gemini 3.8 Flash.
OpenAI, the company behind GPT models, had only two models in the top 20. GPT-5.6 Sol ranked eighteenth, and GPT-5.5 came in twentieth. Two Chinese-origin models, Kimi K3 and GLM 5.3, also made the top 20. The Arena rankings not only provided a general overview but also assessed models on specific tasks to highlight their strengths. In most of these categories, Anthropic and OpenAI models dominated the rankings.
For web development, GPT-6 Astra, which OpenAI described as its best model for software development, topped the WebDev ranking. It outperformed Claude Fable 5.1 Max and Claude Opus 5 Max. In image analysis, Claude Fable 5 was the most performant model, beating Qwen 3.8 and the high-performance version of Claude Opus 4.7. Meta had three models in the top 10 for this category. In document analysis, seven of the top ten positions were taken by Anthropic models, with Claude Opus 5 leading the way. Three OpenAI models completed the top 10.
For image generation, the flare and sunburst versions of ChatGPT Images 2.5 recently launched by Microsoft took the top spots in the Text-to-image ranking, pushing their predecessor to third place. Microsoft MAI Image 2.6 ranked fourth. In web search, GPT-5.6 Sol maintained a slight lead with 1,257 Elo points, surpassing the search version of Claude Opus 4.6. The rest of the top 10 included GPT-5.5 Search, Claude Fable 5, and Gemini 3.1 Pro.
Arena relies on user feedback to evaluate AI models. The platform organizes matches between two models with hidden identities, using the same prompt for both. Users then choose which model performs better. Each model is assigned an Elo score, which changes based on the outcomes of these comparisons. A model earns points when it beats a higher-ranked competitor and loses points when it loses to a lower-ranked one. This system aims to provide an objective measure of AI performance based on user experience.
Anthropic Models Dominate AI Performance Rankings in September 2026
AI-rewritten from original reportingHow it works
ai-rankingsanthropicclaudeopenaimetagpt



