Top LLMs Misidentify Poisonous Mushrooms in Every Ninth Case, Benchmark Shows
Developer Piotr Migdal tested whether current large language models can be trusted to identify mushrooms from photographs, focusing on the highest-risk step of foraging. The short answer is that relying on them could lead to serious poisoning incidents.
The benchmark used 1040 photographs of 55 mushroom species most common in Poland. All images were taken from the FungiTastic dataset, which itself is built on the Atlas of Danish Fungi. The full underlying collection contains more than 600,000 expert-labeled photographs, and a subset of specimens also carries DNA-based species confirmation.
Each model received one photograph at a time and was instructed to return the five most probable species names in Latin. No fine-tuning or external tools were provided. Gemini 3.8 Flash achieved the highest scores, placing the correct species first in 65 percent of cases and within the top five answers in 85 percent of cases. Gemini 3.6 Flash followed with 64 percent and 85 percent, while Gemini 3.7 Flash recorded 61 percent and 82 percent.
At the bottom of the accuracy table were Qwen 3.8 27B, MiniMax-M3, and Qwen 3.8 Flash. Accuracy figures dropped sharply when safety was considered. Both Gemini 3.8 Flash and Gemini 3.6 Flash classified poisonous mushrooms as edible species in approximately 11 percent of cases. The error rate rose to 24 percent for GPT-5.6 Sol, 29 percent for Claude Opus 5, and 36 percent for Qwen 3.8 27B.
Models were never asked directly whether a find was safe to eat. Instead, the researcher mapped the returned species names against a separate table of edibility. Some models refused to provide any classification at all for certain images.
Related articles
HTTP Methods Explained: GET, POST, PUT, PATCH, DELETE and the New QUERY Standard
HTTP methods define the actions a client requests from a server regarding a resource. The core semantics are outlined in RFC 9110, with extensions for specialized protocols. A new standardized method called QUERY was introduced in June 2026 via RFC 10008 to handle complex queries that include a request body while remaining safe and idempotent. The article details safe and idempotent properties, compares each method including GET, HEAD, POST, PUT, PATCH, DELETE, OPTIONS, TRACE, CONNECT, and QUERY, and explains their correct usage to avoid breaking caches, proxies, and infrastructure expectations. It also covers WebDAV extensions and other registered methods in the IANA registry.
From Web Perimeter Breaches to Domain Takeover: How Standoff Hackbase Trains Pentesters on Real Corporate Infrastructure
wr3dmast3r, a senior pentester and BSCP certification guide author, rose to first place on the Standoff Hackbase ranking by shifting focus from initial perimeter access to full internal infrastructure compromise. The platform replicates large-scale corporate networks from various industries, forcing participants to map service relationships, harvest credentials, escalate privileges, and chain pivots across segments. Unlike CTF challenges that end with a single flag, Hackbase tasks require building complete attack paths that can lead to data theft, process disruption, or cross-domain movement. The interview highlights practical techniques such as time-boxing hypotheses, manually modeling infrastructure after automated scans, and using AI only as an information accelerator rather than an autonomous operator. wr3dmast3r also details a memorable chain that began with a bot, moved through VPN and Outlook access, leveraged SCCM tokens for privilege escalation, and ended with compromise of a second domain containing the target system.
OTUS Publishes September Digest of Free Lessons on Linux Administration, PostgreSQL, CI/CD and Infrastructure Security
OTUS has released a new digest listing free September webinars aimed at infrastructure engineers, DevOps specialists and system administrators. The program covers practical topics including Linux server configuration, PostgreSQL 18 performance tuning, high-availability clusters with Patroni, CI/CD pipelines in GitLab, eBPF observability and infrastructure security practices. All sessions are delivered by practicing OTUS instructors who share real-world production experience. Separate tracks address RAID and LVM management, GPO policies, release management in 1C environments, Go profiling, mitmproxy traffic analysis and responsible use of AI tools for incident investigation and code review. The webinars run throughout September at 19:00 or 20:00 Moscow time and require only free registration. The digest also includes sessions on career growth from tech lead to CTO and effective responsibility distribution for team leads.
September 2026 AI Model Rankings: Fable 5.1 Tops Intelligence Index as Competition Tightens Across GPT-5.6 Sol, Grok 4.6 and Muse Spark 1.3
The beginning of September 2026 marked a rare moment when the list of top language models had to be almost entirely rewritten. Anthropic released Fable 5.1 and the limited Mythos 5.1, while Meta updated Muse Spark to version 1.3, Google introduced Gemini 3.8 Flash, and Alibaba refreshed Qwen3.8-Max. Existing models including GPT-5.6 Sol, Grok 4.6, Kimi K3, GLM-5.3 and DeepSeek V4 Pro remain competitive. Traditional rankings from smartest to least capable have become difficult because modern models operate in multiple reasoning-depth modes where low, high and max settings can differ by ten or more points on the same test. The market is better viewed as several overlapping races where Fable 5.1 leads in complex reasoning quality, GPT-5.6 Sol and Grok 4.6 deliver near-top performance at lower cost, and Muse Spark 1.3 excels in price-performance. Independent Artificial Analysis Intelligence Index scores, context windows, API pricing and tool-use capabilities now determine practical choices more than raw benchmark numbers.