Automating Malware Reverse Engineering with Local LLMs, PyGhidra and Neo4j Graphs
A researcher has built an automated system that combines PyGhidra, Neo4j and a local large language model to speed up the initial analysis of malicious binaries. The project addresses the main drawbacks of sending decompiled code to cloud LLMs: leakage of indicators of compromise, possible censorship of sensitive modules, and rapid context overflow when a sample contains 100–200 functions.
Architecture and data flow
The pipeline begins with ghidra_extractor.py, which uses PyGhidra to disassemble and decompile a binary without launching the Ghidra GUI. For every function the extractor records its address, name, size, cyclomatic complexity, decompiled C pseudocode (truncated at 6000 characters with a 30-second timeout), call-graph neighbors, string references and imported API calls. Metadata such as PE format, architecture and all standard file hashes are also stored. All extracted entities are loaded into Neo4j as distinct node types: Sample, Function, String, Import, Analysis, Capability and Behavior.
Next, worker.py and mcp.py iterate over functions and send each one to a locally hosted Qwen3 model running inside LM Studio. The prompt includes both the decompiled code and the function’s position in the call graph, allowing the model to understand caller-callee relationships. The model returns structured JSON describing the function’s purpose, extracted IOCs, tags, network indicators, command-line usage and anti-analysis techniques.
aggregator.py then aggregates tags across all functions into Capability nodes that carry a strength weight reflecting how often each capability appears. builder.py combines these capabilities with the call graph to detect higher-level behavioral patterns such as file encryption combined with network activity. Finally, generator.py produces a human-readable report by querying the enriched graph.
Graph data model
Storing functions in a flat table makes it difficult to answer questions such as whether five anti-debug functions call one another. In the Neo4j model a Function node is connected to other Function nodes via CALLS edges, to String nodes via REFERENCES edges, and to Import nodes via USES edges. Capability and Behavior nodes sit above the function level and are linked through weighted relationships, enabling concise Cypher queries instead of recursive joins.
Validation on WannaCry
The pipeline was tested on a WannaCry sample obtained from MalwareBazaar. Processing 195 functions took 126 minutes on a machine with 60 GB RAM and an Nvidia 5070 Ti. The generated report identified three behavioral patterns with 100 percent confidence: File Encryption (supported by 40 functions), C2 Communication and Anti-Analysis. A publicly available YARA rule confirmed the anti-debug findings discovered by the system.
The complete source code is available in the repository Dvoranchik/MLMalwareAnalyzer. Future work will focus on refining call-graph traversal and prompt engineering to surface additional indicators.
Related articles
Cybercriminals Distribute NJRAT, DCRAT and Chaos via Fake GTA 6 Downloads
Threat actors are leveraging anticipation around GTA 6 to spread multiple malware families through fake game downloads. Security researchers at Huntress identified campaigns that combine search-engine poisoning, gaming forums, torrent sites and social-media posts to deliver oversized fake ISO files exceeding 100 GB. Victims who execute the installer see Russian-language messages claiming an invalid crack or missing license, while remote-access tools and data stealers run silently in the background. The delivered payloads include NJRAT and DCRAT for keystroke logging, screen capture and webcam access, Mercurial Grabber for harvesting browser credentials and Discord tokens, and the Chaos wiper that encrypts small files and overwrites larger ones. The operation primarily targets Russian-speaking gamers, as indicated by the ransom note and error messages. Analysts recommend avoiding unofficial downloads and isolating any compromised systems immediately.
Smartphone Spyware: How Devices Collect and Exfiltrate Data Even Without Internet Access
Modern smartphones continue gathering sensor data including microphone, camera, gyroscope, and satellite navigation even when Wi-Fi and mobile data are disabled. Information is stored locally and transmitted only when connectivity is restored. The Find My Device feature from Google and Find My from Apple allow location reporting for hours after the device is powered off via a separate Bluetooth chip. In August 2026, ThreatFabric disclosed the Manic trojan that uses Wi-Fi Direct, Bluetooth RFCOMM, and BLE GATT to relay encrypted data through nearby infected devices when direct internet access is unavailable. The malware supports multi-hop routing of up to four intermediate devices. Everyday users face greater risk from over-privileged applications and pre-installed malware on gray-market devices than from sophisticated offline exfiltration techniques. A detailed checklist covers purchase hygiene, permission audits, and recovery steps after suspected compromise.
Reconstructed Stuxnet Source Code Published on GitHub with Build Instructions
An unknown researcher has released a reconstructed version of the Stuxnet worm source code on GitHub, including reverse-engineering results and assembly instructions. Stuxnet was discovered in 2010 and is widely attributed to the joint US-Israeli Olympic Games operation targeting Iran's Natanz uranium enrichment facility. The malware specifically attacked Siemens industrial controllers by altering frequency converter operations to physically damage centrifuge rotors while falsifying operator displays. Propagation relied on USB drives, network shares, and a Windows Print Spooler vulnerability, combined with stolen Realtek and JMicron driver-signing certificates. The worm also injected itself into Siemens WinCC and Step 7 software to intercept communications with programmable logic controllers. Due to a flaw in its environment checks, Stuxnet escaped the target network and spread publicly before its built-in June 2012 self-destruct date. Researchers are advised to analyze the code only inside fully isolated virtual machines without network access.
Researchers Create 0-Click WeChat Worm That Hijacks Accounts via Incoming Calls
Security researchers at Calif have developed a 0-click worm capable of compromising WeChat accounts through incoming voice calls without any user interaction. The exploit requires only that the attacker already exists in the victim's contact list and works across Android and iOS devices. In demonstrations, an infected Android device called an iPhone to seize control of its WeChat account while the call continued ringing, after which the compromised iPhone targeted another Android device. The worm spreads automatically between trusted contacts, functioning like a classic network worm but using WeChat profiles as propagation nodes. Even if the victim answers or declines the call, the exploit can persist or be retried, for example during nighttime hours. After account takeover, attackers gain full control to read and send messages, make calls, and impersonate the owner, while the underlying smartphone itself remains unaffected. Tencent received notification in July, released patches in versions 8.0.77 for Android and 8.0.76 for iOS, and server-side blocking was confirmed by late August, with no real-world attacks observed so far.