Fully autonomous AI penetration testing agent

PentAGI runs the penetration test while you watch

Describe the objective and the target. PentAGI plans the engagement, drives a real terminal, browser, editor, and search inside a sandboxed Kali environment, and reports every finding with the evidence behind it.

See how it works
Plans from $7.50/moFully autonomous agent200+ Kali toolsSandboxed Docker runtime12+ LLM providers
Engagement console

Scope the target and let the agent work

Start with a sentence. PentAGI handles reconnaissance, exploitation attempts, privilege escalation, and reporting, then hands back a report that links every finding to the command that produced it.

Autonomous engagementSupervised stepsRecon only
@ Reference an asset, a scope file, or a previous engagement to reuse its knowledge graph.
Scope targets (max 20)0 / 20
Add domains, IP ranges, or repositoriesDomains, CIDR ranges, or repositories. Authorization is required for every target you add.
Context files (max 5)0 / 5
Upload scan results, notes, or credentialsNmap XML, Burp exports, markdown notes, or CSV, up to 25 MB each.
Engagement scope

Describe the target in plain language

Give the agent an objective, a target, and any rules of engagement. PentAGI plans the path itself: reconnaissance, exploitation attempts, privilege escalation, and reporting.

Agent toolchain

One agent, terminal and browser

Every step runs inside a sandboxed container with a real terminal, a headless browser, a code editor, and web search, so the agent can act on what it discovers.

  • Terminal and Kali Linux toolchain
  • Browser automation for live targets
  • Editor for payloads and scripts
  • Search for CVEs and public intelligence
Memory and reporting

Findings stay linked to evidence

Commands, outputs, and observations are stored in PostgreSQL and indexed in a knowledge graph, so the final report cites the exact step that produced each finding.

  • Knowledge graph RAG with Graphiti and Neo4j
  • Full command and output history
  • Structured results and actionable insights
  • REST and GraphQL access to every run

Only test systems you are authorised to test. Every engagement requires an explicit target list, runs in an isolated sandbox, and keeps a complete command and output history for audit.

Engagement patterns

What an autonomous run looks like

The same agent handles reconnaissance, web application testing, privilege escalation paths, and regression testing. Each pattern starts from a plain-language objective.

External attack surface

Enumerate a public domain: subdomains, DNS records, open ports, and the version fingerprint of every exposed service.

target: acme.example.com

Web application assessment

Drive a headless browser and a proxy through the app, capture every request, and follow the parameters that look injectable.

scope: staging web app

Privilege escalation paths

Collect credentials from the environment, then map the shortest path from a low-privilege foothold to domain admin.

goal: privilege escalation

Continuous regression testing

Re-run the same engagement after every release and diff the findings against the previous run.

cadence: every deploy

Vulnerability research

Correlate banner versions with CVEs and public exploit write-ups before attempting anything against the target.

source: CVE + vendor docs

Detection validation

Produce realistic adversary activity and check whether your detections and alerts actually fire.

goal: validate detections
Comprehensive autonomous platform

Everything the agent needs to run a real test

PentAGI combines an isolated execution sandbox, a professional toolchain, and persistent memory so long engagements stay coherent and auditable.

Sandboxed by default

The agent runs inside an isolated Docker environment with its own network and filesystem, so testing never touches your host.

Autonomous execution

PentAGI detects the next step from the previous result and performs it, instead of waiting for you to drive every command.

Browser integration

A headless browser fetches live pages, follows redirects, and reads documentation the moment the agent needs it.

Knowledge graph RAG

Semantic memory powered by Graphiti and Neo4j keeps entities, credentials, hosts, and findings connected across the run.

Complete history

Every command, output, and decision is persisted in PostgreSQL, so any terminal step can be replayed and audited later.

Professional Kali toolchain

Over 200 penetration testing tools ship in the optimized Docker images, from nmap and ffuf to Metasploit and hydra.

Self-hosted

Run the whole platform on your own infrastructure and keep engagement data inside your perimeter.

Modern operator UI

A sleek, intuitive console shows the live plan, current task, terminal stream, and accumulated findings side by side.

Langfuse integration

Control and monitor agents while they work: token usage, tool calls, latency, and cost per engagement.

Version control

Tool state and agent configuration are tracked in a Git project, so runs are reproducible.

Customizable API

Drive engagements from your own tooling through the REST and GraphQL APIs.

Internet search

The agent looks up CVEs, exploits, and vendor documentation through Google, Tavily, and Traversaal.

Five-step workflow

How PentAGI runs a penetration test

From objective to report, the agent keeps a plan you can follow and intervene in at any point.

01

Define the task

Describe the penetration testing objective, the target, and any constraints the agent must respect.

02

AI analysis

PentAGI decomposes the objective, picks a strategy, and builds a plan before touching the target.

03

Automated testing

The agent executes the plan in a sandboxed Kali environment, adapting as each result comes back.

04

Real-time monitoring

Watch the plan, terminal stream, and findings update live, and intervene whenever you want.

05

Results and reporting

Get a structured report that links every finding to the evidence that produced it.

Example objective: Assess the external perimeter of acme.example.com. Enumerate subdomains and exposed services, prioritise anything internet-facing with an outdated component, and attempt to demonstrate impact without disrupting production.
AI service providers

12+ LLM providers for flexible testing

Bring your own key or run entirely locally. PentAGI supports the major model providers plus any OpenAI-compatible endpoint through LiteLLM.

OpenAI

GPT-5+ reasoning models for complex security analysis.

Anthropic

Claude 4+ series with exceptional reasoning capabilities.

Google Gemini

Multimodal models with advanced thinking.

AWS Bedrock

Enterprise-grade foundation models.

Ollama

Local inference for zero-cost private testing.

Custom

Any OpenAI-compatible endpoint.

DeepInfra

Fast inference for cost-effective operations.

OpenRouter

Newest models through one unified API.

DeepSeek

Reasoning models for complex problem solving.

GLM (Zhipu AI)

Models with strong multilingual capabilities.

Kimi (Moonshot)

Long-context models up to 200k tokens.

Qwen (Alibaba)

Open-source models with strong reasoning.

Observability stack

Monitor every agent decision

PentAGI ships with a complete observability stack so you can analyse what the agent did, why it did it, and what it cost.

Grafana

Visualize metrics and logs from every engagement.

Loki

Log aggregation for agent and tool output.

ClickHouse

Column-oriented analytics database for run history.

Jaeger

Distributed tracing across agent steps.

OpenTelemetry

Vendor-neutral instrumentation for the whole stack.

VictoriaMetrics

High-performance time series metrics.

Real testing workflows

Who runs PentAGI engagements

Security teams use PentAGI for the work that is repetitive, time-boxed, or needs to run on every release.

Offensive security teams

Red teams

Run repeatable internal engagements and keep the human operator on the decisions that matter.

Detection engineering

Blue teams

Generate realistic adversary activity to validate detections and alerts end to end.

Independent researchers

Bug bounty hunters

Reconnaissance, asset discovery, and repetitive probing handled by the agent while you focus on exploitation logic.

Product security

AppSec teams

Drive web application assessments with browser automation, proxies, and evidence capture.

Audit and assurance

Compliance

Produce auditable evidence trails for each test through the persisted command history.

Labs and CTFs

Security research

Automate the tedious parts of a lab environment and study how an autonomous agent reasons.

Plans for every engagement

Start monthly, or save 17% when you pay annually. Every plan unlocks the same PentAGI workflow - describe an objective and the agent plans, executes, and reports the engagement inside a sandboxed environment - so you only choose how much testing capacity you need.

Starter

$9
$7.50 /month
A full year of PentAGI engagements at the lowest monthly price.
Pay annually and save 17%.

Included

  • 1,200 credits a year (100 a month)
  • Autonomous planning and execution
  • Sandboxed Docker runtime with 200+ Kali tools
  • Browser, editor, and web search tools
  • Sign in with Google
  • Secure checkout with Stripe
MOST POPULAR

Pro

$19
$15.83 /month
A full year of weekly engagements with the Pro toolchain.
Best value once PentAGI is part of your normal weekly rhythm.

Included

  • Everything in Starter
  • 2,400 credits a year (200 a month)
  • Knowledge graph RAG with Graphiti and Neo4j
  • Full command history and evidence trails
  • Structured findings and reporting
  • Secure checkout with Stripe
HIGHEST CAPACITY

Premium

$49
$40.83 /month
For heavy, continuous testing when you re-run engagements often.
Choose Premium when PentAGI stays in your operating loop all year.

Included

  • Everything in Pro
  • 3,600 credits a year (300 a month)
  • Highest annual testing capacity
  • Best for continuous regression testing after every release
  • Priority support
  • Secure checkout with Stripe
Frequently asked questions

PentAGI FAQ

Answers about how the agent runs engagements, which models are supported, and how credits are consumed.

What is PentAGI?+

PentAGI is a fully autonomous AI agent for complicated penetration testing tasks. It plans an engagement, executes it with a terminal, browser, editor, and search inside a sandboxed environment, and produces a structured report.

How do I use PentAGI online?+

Sign in with Google, pick a plan, describe the objective and target in the console, and start the engagement. Nothing needs to be installed to try the hosted workspace.

Does the agent run commands on my own machine?+

No. Every command runs inside an isolated Docker sandbox with its own network and filesystem, and all output is stored for the run.

Which AI models can I use?+

More than 12 providers are supported, including OpenAI, Anthropic, Google Gemini, AWS Bedrock, Ollama, DeepInfra, OpenRouter, DeepSeek, GLM, Kimi, and Qwen, plus any OpenAI-compatible endpoint or a LiteLLM proxy.

How does the agent remember what it found?+

Commands and outputs are persisted in PostgreSQL, and semantic memory is kept in a knowledge graph powered by Graphiti and Neo4j, so findings stay connected across a long engagement.

Can I watch the engagement while it runs?+

Yes. The console streams the plan, the active task, the terminal output, and new findings in real time, and monitoring integrations such as Langfuse, Grafana, and Jaeger are supported.

Do I need my own infrastructure?+

No. The hosted workspace provisions an isolated environment per engagement. Self-hosting on your own infrastructure is also supported if engagement data must stay in your perimeter.

How are credits used?+

Each engagement consumes credits based on the depth of the run, the number of tools invoked, and the reasoning model selected. The console shows an estimate before you start.

Is pentagi.homes the official PentAGI website?+

No. pentagi.homes is an independent third-party workspace and is not affiliated with, endorsed by, sponsored by, or operated by VXControl L.L.C-FZ or any of its affiliates.

Ready to enhance your security?

Start using PentAGI today and experience AI-driven penetration testing on your own targets.

Open the console
ShipAny Template Two