安全客July 28, 2026🇨🇳Translated from Chinese

PentesterFlow Launches Open-Source AI CLI Tool for Penetration Testers and Bug Bounty Hunters

PentesterFlow is a new open-source, human-in-the-loop AI command-line tool built for penetration testers and bug bounty hunters. It automates the entire workflow from information gathering to report generation without removing analyst control.

Most agentic AI security tools suffer from hallucinated findings, poor context retention, and weak tool integration. PentesterFlow tackles these problems with built-in pentesting skills, evidence-based vulnerability confirmation, and continuous local learning capabilities.

The tool connects to local or hosted large language models such as Ollama, LM Studio, Kimi, Groq, Gemini, DeepSeek, and OpenRouter. It plans actions against defined targets, runs real security testing tools, and always requests explicit analyst approval before executing sensitive commands.

Core Capabilities

PentesterFlow covers the full penetration testing lifecycle: scoping, reconnaissance, enumeration, vulnerability validation, coverage tracking, reporting, and ongoing learning.

  • Model backends: Ollama, LM Studio, Kimi, Groq, Gemini, DeepSeek, OpenRouter and OpenAI-compatible APIs
  • Built-in skills: Recon, Web vulnerabilities, SSRF, SSTI, JWT, GraphQL, race conditions, subdomain takeover, Supabase, deserialization
  • Toolset: Shell/Bash, HTTP, Burp Suite bridge, browser capture, MCP support, file operations, grep/glob
  • Report output: Confirmed findings with PoCs, impact analysis, remediation advice, and ready-to-use curl commands

In a live demonstration, the tool loaded a web vulnerability skill module, sent HTTP requests against an order API, and automatically confirmed a high-severity IDOR vulnerability, writing the evidence-backed finding directly to a Markdown file.

Security and Integration

PentesterFlow enforces permission-based tool calls, blocks dangerous shell command patterns, and redacts credentials during compression and snapshot operations. It also offers a YOLO mode that auto-approves all actions inside isolated lab environments.

The tool integrates directly with Burp Suite through a companion bridge, allowing testers to import captured traffic into the CLI and export confirmed findings back as Burp issues.

Installation and Usage

Installation uses a simple shell script on macOS/Linux or PowerShell on Windows to fetch the latest standalone binary and verify its SHA-256 checksum. Users can pin specific versions, select local models such as Ollama’s qwen2.5-coder, or connect to hosted providers, then set targets with the /target command and issue natural-language instructions.

Analysts working on sensitive targets are reminded that PentesterFlow is intended solely for authorized security testing, as approved commands can execute real shell operations and send live HTTP requests.

Industry Positioning

Alongside projects such as PentAGI and PentestGPT, PentesterFlow distinguishes itself through transparent, reproducible evidence chains and mandatory analyst approval, making it suitable for security teams cautious about fully autonomous penetration testing agents.

Related articles

HabrAI Security

Agent-Ops 0.4.0 Released: Methodology for Secure Human-AI Collaboration in IT Operations

Sergey Zhitinsky, founder of Git in Sky, has published the public normative candidate for Agent-Ops 0.4.0, an open industry methodology governing how engineers and AI agents jointly handle IT infrastructure tasks. The framework keeps humans firmly in the decision-making loop while using deterministic programs for data collection and approved changes. It addresses risks such as prompt injection through processed data, unverified model outputs, and unclear accountability when AI recommendations lead to incidents. The methodology divides work across eight explicit steps and three separate planes: data, governance, and independent verification performed by a Guardian role. Two additional companies have joined as maintainers following agreements at the IT Elements 2026 conference, turning the project into a multi-organization effort. Contributors are invited to help refine contracts, schemas, and operational scenarios through GitHub and GitVerse.

HabrAI Security

ProxyKey MCP: Securing API Access for AI Agents Without Exposing Credentials

ProxyKey has released an MCP server that allows AI coding agents such as Claude Code and Cursor to manage API credentials without ever reading the actual secret values. The solution addresses the risk that any key visible to an agent becomes compromised through logging, tracing, or prompt injection. Real provider keys are stored encrypted with AES-256-GCM and never returned by any API endpoint after initial entry. Agents instead receive limited virtual passes that support IP binding, rate limits, TTL, and detailed request logging. A pending-secret workflow lets agents prepare services before the real token exists, with the human entering the secret only through a web panel. The approach deliberately restricts the MCP tool contract so no operation can read or return secret values.

HabrAI Security

Shadow AI in CI/CD: Why AI Agents Must Be Modeled as Security Threats

A new analysis from the CNCF highlights the growing risks of Shadow AI within continuous integration and continuous deployment pipelines. The report argues that AI agents should be treated as potential threats rather than simple productivity tools. Starting from a developer's laptop and extending to Kubernetes clusters, these agents can introduce unauthorized access paths and data exposure risks. Security teams are urged to incorporate AI agent behavior into formal threat modeling exercises. The discussion emphasizes the need for visibility and control over autonomous AI components operating in production environments.

HabrAI Security

Detecting Lateral Movement with Neural Networks Trained Solely on Synthetic Data

A researcher generated entire corporate network histories using a 135-line configuration file to create synthetic authentication logs containing lateral movement attacks. Neural networks trained exclusively on these artificial datasets were then evaluated against 1.65 billion real authentication events from Los Alamos National Laboratory, including 749 red team events across 301 compromised machines. The best ensemble of six models flagged 3.6 million hourly machine windows and placed 16 genuine attacks among the top 23 highest-scoring entries, producing only seven false positives. In comparison, a simple threshold counter required 161,000 false alarms to reach the same detection level. The approach also demonstrated an iterative feedback loop where detector errors directly informed refinements to the synthetic world generator. The work shows that synthetic data can reach AUC performance comparable to models trained on real labeled attacks while providing full control over the underlying attack definitions.