SecuritylabSeptember 4, 2026🇷🇺Translated from Russian

September 2026 AI Model Rankings: Fable 5.1 Tops Intelligence Index as Competition Tightens Across GPT-5.6 Sol, Grok 4.6 and Muse Spark 1.3

The start of September 2026 proved to be a rare moment when the roster of the strongest language models had to be almost completely rewritten. Anthropic introduced Fable 5.1 and the restricted Mythos 5.1, Meta updated Muse Spark to version 1.3, Google launched Gemini 3.8 Flash, and Alibaba refreshed Qwen3.8-Max. Summer releases such as GPT-5.6 Sol, Grok 4.6, Kimi K3, GLM-5.3 and DeepSeek V4 Pro stayed in the race and continue to compete with the newcomers.

Compiling a conventional ranking from smartest to least capable has become difficult. Contemporary models run in several reasoning-depth modes, and the gap between low, high and max settings can exceed ten points on a single test. Speed, token consumption, long-context pricing, image support, tool-use quality and the ability to execute long autonomous action sequences now vary widely.

It is therefore more useful to treat the market as several overlapping races. Fable 5.1 currently claims the highest quality for complex reasoning. GPT-5.6 Sol and Grok 4.6 deliver comparable performance at lower cost. Muse Spark 1.3 unexpectedly entered the top tier on capability-to-cost ratio. Gemini 3.8 Flash bets on speed and multimodality. Kimi, GLM, Qwen and DeepSeek have nearly erased the former boundary between Chinese and American models.

How to Compare Models Without Misleading Yourself

The Artificial Analysis Intelligence Index serves as a convenient common scale. This independent metric combines tests on knowledge, programming, scientific reasoning, tool use and long-horizon agent tasks. The methodology is more reliable than vendor tables because models are evaluated under comparable conditions.

Scores cannot be read as an intelligence quotient. A result of 66 versus 61 does not mean the first model is eight percent smarter. The index reflects average performance on a specific test suite. On Russian-language article writing, accounting-document analysis, website layout or debugging enterprise applications the ordering may differ.

Reasoning modes alter the picture even more. GPT-5.6 Sol scores around 61 in max mode, 59 in xhigh and 57 in high. Fable 5.1 reaches 66 only at maximum depth. Gemini 3.8 Flash scores 59 in high, 57 in medium and 52 in low. Comparing Gemini against Claude without specifying the mode is about as useful as comparing cars without stating engine power.

Market Leaders at the Start of September 2026

The current top tier is unusually dense. After Fable 5.1 the gap between several flagships fits within a few points. Expensive models do not always complete a given task more cheaply; some systems generate far more intermediate text, reason longer or invoke external tools more frequently.

  • Claude Fable 5.1, max – 66 points, 1 M context, $10/$50 per million tokens, available
  • Claude Opus 5, max – 63 points, 1 M context, $5/$25, available
  • Muse Spark 1.3, max – 62 points, 1 M context, price not announced, limited access
  • GPT-5.6 Sol, max – 61 points, ~1.05 M context, $4/$20, available
  • Grok 4.6, high – 61 points, 500 k context, $2/$6, available
  • Muse Spark 1.3, xhigh – 61 points, 1 M context, $1.25/$4.25, available
  • Kimi K3, max – 60 points, ~1 M context, $3/$15, open weights
  • GLM-5.3, max – 60 points, up to 1 M context, $1.40/$4.40, open weights
  • Gemini 3.8 Flash, high – 59 points, ~1.05 M context, $0.75/$3.75, available
  • Qwen3.8-Max – 58 points (previous snapshot), 1 M context, ~$2/$6, updated to 0902
  • DeepSeek V4 Pro 0813, max – 53 points, 1 M context, $0.66–1.32/$1.98–3.96, open weights

Additional caveats apply. Independent tests of the updated Qwen3.8-Max-0902 have not yet produced a comparable final score, so the 58-point figure refers to the prior version. Muse Spark 1.3 maximum mode remains in limited testing; the public xhigh mode scores 61. For Grok 4.6 the best independent result occurs in high mode, while xhigh unexpectedly scores slightly lower.

Pricing also requires notes. The Gemini 3.8 Flash tariff is temporarily reduced until the end of 2026. GPT-5.6 Sol is sold at a temporarily lowered price. DeepSeek pricing changes between peak and off-peak hours. Grok 4.6 doubles in price for requests longer than 200 k tokens. A simple API-price column therefore hides half of the real economics.

Related articles

SecuritylabOther

HTTP Methods Explained: GET, POST, PUT, PATCH, DELETE and the New QUERY Standard

HTTP methods define the actions a client requests from a server regarding a resource. The core semantics are outlined in RFC 9110, with extensions for specialized protocols. A new standardized method called QUERY was introduced in June 2026 via RFC 10008 to handle complex queries that include a request body while remaining safe and idempotent. The article details safe and idempotent properties, compares each method including GET, HEAD, POST, PUT, PATCH, DELETE, OPTIONS, TRACE, CONNECT, and QUERY, and explains their correct usage to avoid breaking caches, proxies, and infrastructure expectations. It also covers WebDAV extensions and other registered methods in the IANA registry.

SecuritylabOther

From Web Perimeter Breaches to Domain Takeover: How Standoff Hackbase Trains Pentesters on Real Corporate Infrastructure

wr3dmast3r, a senior pentester and BSCP certification guide author, rose to first place on the Standoff Hackbase ranking by shifting focus from initial perimeter access to full internal infrastructure compromise. The platform replicates large-scale corporate networks from various industries, forcing participants to map service relationships, harvest credentials, escalate privileges, and chain pivots across segments. Unlike CTF challenges that end with a single flag, Hackbase tasks require building complete attack paths that can lead to data theft, process disruption, or cross-domain movement. The interview highlights practical techniques such as time-boxing hypotheses, manually modeling infrastructure after automated scans, and using AI only as an information accelerator rather than an autonomous operator. wr3dmast3r also details a memorable chain that began with a bot, moved through VPN and Outlook access, leveraged SCCM tokens for privilege escalation, and ended with compromise of a second domain containing the target system.

HabrOther

OTUS Publishes September Digest of Free Lessons on Linux Administration, PostgreSQL, CI/CD and Infrastructure Security

OTUS has released a new digest listing free September webinars aimed at infrastructure engineers, DevOps specialists and system administrators. The program covers practical topics including Linux server configuration, PostgreSQL 18 performance tuning, high-availability clusters with Patroni, CI/CD pipelines in GitLab, eBPF observability and infrastructure security practices. All sessions are delivered by practicing OTUS instructors who share real-world production experience. Separate tracks address RAID and LVM management, GPO policies, release management in 1C environments, Go profiling, mitmproxy traffic analysis and responsible use of AI tools for incident investigation and code review. The webinars run throughout September at 19:00 or 20:00 Moscow time and require only free registration. The digest also includes sessions on career growth from tech lead to CTO and effective responsibility distribution for team leads.

AntiMalwareOther

Top LLMs Misidentify Poisonous Mushrooms in Every Ninth Case, Benchmark Shows

Polish developer Piotr Migdal evaluated leading large language models on their ability to identify mushrooms from photographs, using a dataset of 1040 images covering 55 species common in Poland. The images came from the FungiTastic dataset derived from the Atlas of Danish Fungi, with expert labels and partial DNA confirmation. Models were asked to return the five most likely species names in Latin without additional training or tools. Gemini 3.8 Flash performed best with 65 percent top-1 accuracy and 85 percent top-5 accuracy, followed closely by other Gemini variants. However, safety-critical errors remained high: Gemini models labeled poisonous mushrooms as edible in roughly 11 percent of cases, while GPT-5.6 Sol reached 24 percent, Claude Opus 5 reached 29 percent, and Qwen 3.8 27B reached 36 percent. The study did not ask models directly whether a mushroom was edible; species identifications were later cross-checked against toxicity tables.