elder-plinius/T3MP3STPublic

autonomous red teaming platform; multi-agent offensive-security meta-harness

AI summary: A multi-agent offensive-security framework that equips AI coding assistants for autonomous red teaming.

Stars
6.3K
+25 today
Forks
1.3K
Watchers
65
Open issues
4
Open PRs
3
Contributors
~37
Commits
204
Branches
2

TypeScriptAGPL-3.0Created Jul 2, 2026Last push 26d agoLatest release v1.0.0+71 stars this week+375 this month

Quick answers

What is T3MP3ST?
A multi-agent offensive-security framework that equips AI coding assistants for autonomous red teaming.
What does T3MP3ST do?
T3MP3ST is an advanced offensive-security meta-harness designed to transform standard AI coding agents into autonomous vulnerability hunters. It wraps around existing agents like Claude Code or local offline models, providing them with the necessary tools to execute the complete kill chain: reconnaissance, exploitation, and reporting. Operating locally without requiring new cloud tenants or API keys, it allows security teams to conduct keyless, self-hosted warfare against authorized targets. It boasts high success rates on vulnerability benchmarks and can execute zero-day hunts directly from a browser War Room or the CLI.
Who is T3MP3ST for?
Designed strictly for authorized cybersecurity professionals, penetration testers, and red team operators. It requires deep knowledge of offensive security methodologies and a controlled, authorized environment for deployment.
How do I get started with T3MP3ST?
git clone https://github.com/elder-plinius/T3MP3ST.git
How popular is T3MP3ST on GitHub?
elder-plinius/T3MP3ST has 6,299 stars and 1,314 forks on GitHub, and gained 71 stars in the last 7 days.
What license does T3MP3ST use?
elder-plinius/T3MP3ST is released under the AGPL-3.0 license.

Star history

since Jul 29, 2026
02K4K6KJul 2026Aug 2026Sep 2026Oct 2026
6.3K stars as of Oct 2, 2026. Measured daily since Jul 29, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
OctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 10 commits2026-07-03: 11 commits2026-07-04: 2 commits2026-07-05: 21 commits2026-07-06: 15 commits2026-07-07: 2 commits2026-07-08: 4 commits2026-07-09: 2 commits2026-07-10: 5 commits2026-07-11: 4 commits2026-07-12: 1 commit2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 14 commits2026-07-16: 10 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 9 commits2026-07-21: 2 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 1 commit2026-07-27: 1 commit2026-07-28: 1 commit2026-07-29: 0 commits2026-07-30: 2 commits2026-07-31: 3 commits2026-08-01: 4 commits2026-08-02: 1 commit2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits2026-08-09: 0 commits2026-08-10: 0 commits2026-08-11: 0 commits2026-08-12: 1 commit2026-08-13: 0 commits2026-08-14: 0 commits2026-08-15: 0 commits2026-08-16: 0 commits2026-08-17: 0 commits2026-08-18: 0 commits2026-08-19: 0 commits2026-08-20: 0 commits2026-08-21: 0 commits2026-08-22: 0 commits2026-08-23: 4 commits2026-08-24: 1 commit2026-08-25: 0 commits2026-08-26: 0 commits2026-08-27: 0 commits2026-08-28: 0 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 0 commits2026-09-01: 0 commits2026-09-02: 3 commits2026-09-03: 18 commits2026-09-04: 6 commits2026-09-05: 0 commits2026-09-06: 0 commits2026-09-07: 2 commits2026-09-08: 0 commits2026-09-09: 0 commits2026-09-10: 0 commits2026-09-11: 0 commits2026-09-12: 0 commits2026-09-13: 0 commits2026-09-14: 0 commits2026-09-15: 0 commits2026-09-16: 0 commits2026-09-17: 0 commits2026-09-18: 0 commits2026-09-19: 0 commits2026-09-20: 0 commits2026-09-21: 0 commits2026-09-22: 0 commits2026-09-23: 0 commits2026-09-24: 0 commits2026-09-25: 0 commits2026-09-26: 0 commits2026-09-27: 0 commits2026-09-28: 0 commits2026-09-29: 0 commits2026-09-30: 0 commits2026-10-01: 0 commits2026-10-02: 0 commits2026-10-03: 0 commits
160 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Breakout launch

    6,299 stars in 94 days

  • Outside contributions

    80% of recent commits from the community

  • Well documented

    High community health score

What T3MP3ST does

T3MP3ST is an advanced offensive-security meta-harness designed to transform standard AI coding agents into autonomous vulnerability hunters. It wraps around existing agents like Claude Code or local offline models, providing them with the necessary tools to execute the complete kill chain: reconnaissance, exploitation, and reporting. Operating locally without requiring new cloud tenants or API keys, it allows security teams to conduct keyless, self-hosted warfare against authorized targets. It boasts high success rates on vulnerability benchmarks and can execute zero-day hunts directly from a browser War Room or the CLI.

Designed strictly for authorized cybersecurity professionals, penetration testers, and red team operators. It requires deep knowledge of offensive security methodologies and a controlled, authorized environment for deployment.

  • Agent integration: Seamlessly wraps around existing AI coding agents or local models to repurpose them for offensive security.
  • Autonomous kill chain: Automatically executes the full attack lifecycle from initial reconnaissance to exploitation and final reporting.
  • Local, keyless execution: Operates entirely locally without requiring secondary cloud subscriptions or external API keys.
  • War Room interface: Provides both a browser-based dashboard and a CLI for operators to command and monitor the agent swarm.
  • High benchmark performance: Achieves over 90% success rates on complex cybersecurity challenge suites and captures flags autonomously.

Where teams use it

Autonomous red teaming

Security teams deploy the framework against their own infrastructure to autonomously discover and exploit vulnerabilities before attackers do.

Zero-day vulnerability hunting

Researchers use the agent swarm to actively hunt for undocumented vulnerabilities in real-world, post-cutoff software.

Automated penetration testing

Consultants run the harness to quickly map attack surfaces and generate comprehensive exploit reports for clients.

Capture the Flag (CTF) automation

Competitors utilize the framework to autonomously solve complex security challenges and capture flags without human hints.

Getting started: git clone https://github.com/elder-plinius/T3MP3ST.git

README

main branch

🌩️ T3MP3ST 🌩️

 ▄▄▄█████▓▓█████  ███▄ ▄███▓ ██▓███  ▓█████   ██████ ▄▄▄█████▓
 ▓  ██▒ ▓▒▓█   ▀ ▓██▒▀█▀ ██▒▓██░  ██▒▓█   ▀ ▒██    ▒ ▓  ██▒ ▓▒
 ▒ ▓██░ ▒░▒███   ▓██    ▓██░▓██░ ██▓▒▒███   ░ ▓██▄   ▒ ▓██░ ▒░
 ░ ▓██▓ ░ ▒▓█  ▄ ▒██    ▒██ ▒██▄█▓▒ ▒▒▓█  ▄   ▒   ██▒░ ▓██▓ ░
   ▒██▒ ░ ░▒████▒▒██▒   ░██▒▒██▒ ░  ░░▒████▒▒██████▒▒  ▒██▒ ░
   ▒ ░░   ░░ ▒░ ░░ ▒░   ░  ░▒▓▒░ ░  ░░░ ▒░ ░▒ ▒▓▒ ▒ ░  ▒ ░░
     ░     ░ ░  ░░  ░      ░░▒ ░      ░ ░  ░░ ░▒  ░ ░    ░
   ░         ░   ░      ░   ░░          ░   ░  ░  ░    ░
             ░  ░       ░               ░  ░      ░

A multi-agent offensive-security framework, built to turn the AI coding agent you already run into a zero-day hunter.

scores: re-derivable   verify-claims 27/27   PRs welcome   License: AGPL-3.0

Your AI coding agent is already a hacker — T3MP3ST hands it an arsenal.

Point it at an authorized target and the kill chain runs itself: recon → exploit → report, from a browser War Room or the CLI, driven by the agent you're already signed into — Claude Code, Codex, Hermes, OpenCode, Oh My Pi — or a model you run fully offline (Ollama, LM Studio, vLLM). No new API keys, no cloud tenant, no second bill. Your agent is the brain; T3MP3ST is the war machine bolted around it. Self-hosted storm. Keyless warfare. ⚡

And it won't ask you to take its word for it. On XBOW's own 104-challenge suite it scores 90.1% pass@1 — above XBOW's self-reported 85% — alongside hint-free CTF solves and a cold hunt on real, post-cutoff CVEs the model had never seen. Every number in this README recomputes from committed data with one command (npm run verify-claims). Loud about the mission, honest about the build — the status table says exactly what's live, what's scaffolding, and what's still roadmap; full receipts in Benchmarks.

Three things set it apart:

  1. Reproducible. Every number in this README recomputes from committed data — npm run verify-claims re-derives all of them, 27/27 green. A claim that can't be reproduced doesn't ship. No trust-me numbers, ever.
  2. Keyless. The AI coding agent already on your machine is the backbone. No API keys, no second bill, no gatekeeper.
  3. Honest about scope. The status table marks exactly what's stable, experimental, or roadmap — because red-teaming shouldn't be a priesthood, and it damn sure shouldn't run on vibes.

Jump to → Quick start · Usage guide · Developer guide · Updating · What it hunts · What ships today · Benchmarks · Architecture · Docs

⚠️ Authorized use only

T3MP3ST is an offensive security tool, built for authorized testing, research, and education. Point it only at systems you own or have explicit, written permission to test. Unauthorized access to computers, networks, or data is illegal in most jurisdictions — you alone are responsible for how you use this software and for staying inside the law and your rules of engagement. Bring the storm to your targets, not someone else's.

T3MP3ST is provided as-is under the AGPL-3.0 license, with no warranty and no liability for any damage, loss, or misuse. The authors do not endorse, support, or condone unauthorized activity. Get permission. Stay in scope. Don't be a menace. 🫡

Why it exists

Offensive security sits behind years of practice and expensive tooling. The bet behind T3MP3ST is that a coordinated agent swarm puts real bug-hunting in reach of people who never got the invite, across web apps, CTFs, smart contracts, source code, and embedded/robotics OSS. That is an ambitious bet, and the sections below are careful to separate what already works from what is still a bet.

What it hunts

Domain What it does Status
🕸️ Web apps Black-box, external-attacker recon → exploit (XBEN suite) ✅ Stable
🚩 CTF Hint-free, sandbox-jailed solves (Cybench) ✅ Stable
🤖 Robotics / OT / embedded Coordinated-disclosure pipeline for OSS vuln hunting (OSV + live-PoC + refuter) ✅ Pipeline stable
📂 Source code White-box repo analysis with blind master-builder decomposition ✅ Multi-language ingest (web-tree-sitter)
💰 Smart contracts Damn Vulnerable DeFi ⚠️ reproduction, not novel discovery
☁️ Cloud (IaC) Misconfig-detection benchmark (cloud:bench) + opt-in cloud arsenal (aws/az/gcloud + scoutsuite/cloudfox/pmapper; pacu gated) 🚧 IaC-misconfig scaffolding — live-cloud exploitation not yet benchmarked
📱 Mobile Built-in static analyzer (manifest misconfig + secret/cleartext detection, mobile:bench) + opt-in arsenal (mobsfscan/objection/drozer; frida gated) 🚧 static-detection scaffolding — dynamic exploitation not benchmarked
🔩 Binary / RE Decompiled-output sink detector (unsafe-copy / format-string / cmd-injection / int-overflow, binary:bench) + opt-in arsenal (ghidra/radare2/objdump/checksec/strings; gdb gated) 🚧 static sink-detection scaffolding — solving/pwn not benchmarked

Quick start

Fastest path to a running War Room (keyless, ~2 min to set up; mission time depends on the target):

npm install
npm run server        # War Room → http://127.0.0.1:3333/ui/

In the War Room, open Settings and connect a local agent (Claude Code / Codex / Hermes / OpenCode / Oh My Pi). Then describe a target to Op Admiral in plain English and launch. The agent you connected is the brain. No key required.

Prefer to bring a key? Set one and skip the connect step:

export OPENROUTER_API_KEY=...     # or VENICE_API_KEY / ANTHROPIC_API_KEY / OPENAI_API_KEY
export XAI_API_KEY=...            # Grok Build (grok-build-0.1) — xAI's coding model, native tool-calling
export NOVITA_API_KEY=...         # Novita AI's hosted OpenAI-compatible API

Novita AI is an opt-in third-party hosted provider. Requests send the selected model id, prompts/context, generated output, and request metadata to Novita; your API key authenticates those requests. Review Novita's current privacy policy and terms before using sensitive target data. T3MP3ST does not claim independent assurance for sensitive workloads.

Slow local agents can be given more room with T3MP3ST_LOCAL_AGENT_TIMEOUT_MS for each CLI call, T3MP3ST_TASK_TIMEOUT_MS for mission tasks, and T3MP3ST_GENERAL_TIMEOUT_MS for planning requests. Values are milliseconds.

Claude Code session reuse is disabled by default because a resumed coding-agent session may retain provider-side context and ambient tool authority outside the Arsenal receipt boundary. Operators who explicitly trust their local Claude Code configuration as a separate execution authority may set T3MP3ST_TRUST_CLAUDE_SESSION=1; reuse remains isolated to one task and the choice must not be treated as an Arsenal-enforced sandbox.

Or run it fully offline on your own model — no key, no cloud. Defaults to Ollama; point it at any OpenAI-compatible server (LM Studio, vLLM, llama.cpp):

ollama serve && ollama pull llama3                          # or an OpenAI-compatible server
export TEMPEST_LOCAL_BASE_URL=http://localhost:11434/api    # LM Studio: http://localhost:1234/v1
export TEMPEST_LOCAL_MODEL=llama3
npm run build                                               # only needed from a git clone
npx tempest                                                 # → "Change default provider" → local

Tool-calling works on any local model (it's driven over text), so the Arsenal runs even on models without native function-calling.

Check the numbers for yourself:

npm run verify-claims             # re-derives every headline from committed JSON in bench/

Step-by-step operator usage lives in Getting Started. Library/SDK usage, the HTTP API, and MCP setup live in docs/.

Docker

Run T3MP3ST API server in a container (localhost only, not exposed externally):

cp .env.example .env       # configure API keys
docker compose up -d       # API → http://localhost:3333
docker compose logs -f     # view logs

Security Note: Container binds to 127.0.0.1:3333 - accessible only from localhost, not exposed to network.

Test the API:

curl http://localhost:3333/api/health
curl http://localhost:3333/api/bounty/platforms

Execute commands inside the container:

docker compose exec app npm run verify-claims
docker compose exec app npm run cve:bench

Full deployment guide: docs/DOCKER.md.

Updating from upstream

If you installed from a release tarball or copied the tree instead of tracking git pull, use the built-in updater to sync with github.com/elder-plinius/T3MP3ST without losing local secrets or bench output. It shows a numbered plan, asks y/N before changing anything, then runs npm install.

npm run update          # interactive sync from upstream main
npm run update:dry      # preview only — no git or npm changes
npm run update:hard     # hard reset to upstream/main (still restores protected paths)

Works on Windows (PowerShell), macOS, Linux, and WSL. Requires git and npm on your PATH.

Safety modes

The updater is destructive only when explicitly requested:

Command What it does Local changes
npm run update Default interactive sync git merge upstream/main; protected paths backed up and restored
npm run update:dry Read-only preview No git init, fetch, merge, reset, or npm install; safe on tarball installs
npm run update:hard Opt-in hard reset git reset --hard upstream/main; protected paths still restored

All non-dry-run modes ask y/N before changing anything. Pass --force only in trusted automation. On the first-time path (no commits yet), the updater replaces the working tree with the upstream snapshot, but protected paths are backed up first and restored afterward.

Protected paths (inside the repo)

Before replacing files, the updater backs up anything on disk that matches scripts/update-protected.txt, then restores it after sync. Only paths that exist locally are affected — if you never created them, nothing happens.

Path Why it's protected
.env, .env.* API keys and local env overrides. !.env.example is an exception — the template is not protected so upstream can update it.
.keys.local One-off key paste file loaded by bench scripts (e.g. VENICE_API_KEY) without touching .env.
.keys.bounty.json HackerOne / Bugcrowd / similar platform credentials.
bench/cybench/corpus-stage/ Large cloned Cybench corpus (not redistributed; expensive to re-download).
bench/cybench/service-stage/, bench/cybench/challenges/ Per-run Cybench Docker staging and challenge trees (regenerable, but slow to rebuild).
bench/xbow/stage/, bench/xbow/challenges/ XBOW/XBEN challenge staging (large third-party trees).
bench/wild-hunt/ Cold-hunt findings, PoCs, disclosure drafts, and campaign results — pre-coordination vuln material.
bench/decomposition-results/ White-box decomposition run JSON (may contain unreported analysis).
bench/refusal-frontier/ Refusal-boundary probe artifacts (raw model responses).
bench/nyu/ Staged NYU CTF content from nyu-prep.mjs.
docs/disclosures/ Generated vendor disclosure packages (disclosure-gen output).
reports/ Engagement and hunt reports (persistent output; same tree as Docker volume mounts in deployment setups).
evidence/ PoCs, screenshots, and other finding evidence kept across updates.

Add your own patterns in scripts/update-protected.local.txt (optional local overlay; same glob syntax as the main manifest).

Never touched (outside the repo)

These live outside the project tree — an update never reads or writes them:

  • %APPDATA%\t3mp3st-nodejs\Config\config.json (or the macOS/Linux conf store path) — saved by npm run setup
  • War Room browser localStorage on the War Room origin
  • Local agent auth (~/.codex, %LOCALAPPDATA%\hermes, etc.)

What ships today

The framework is an 8-operator kill chain, and this table won't blow smoke about it. Recon is a live, tool-backed engine — and the teeth are already real: 90.1% pass@1 on XBEN, 8/10 held-out post-cutoff CVEs pinned to exact file/line/CWE, and a coordinated-disclosure pipeline that's live enough to have drafts held for vendor coordination right now. What's not proven is the swarm. Each downstream operator — Exploiter, Infiltrator, Exfiltrator, Ghost — runs the same real, tool-backed ReAct loop as recon (real exploit tools, not stubs), but the headline numbers came from a single agent, not the coordinated 8-operator cell, and end-to-end swarm exploitation is unbenchmarked and still unreliable. The engine is real; the swarm is the part still earning its stripes. Loud where we've earned it, blunt about the rest.

Component Status Notes
Re-derivable measurement (verify-claims) ✅ Stable every headline recomputes from committed artifacts
Recon engine ✅ Stable drives nmap / DNS / HTTP / fingerprinting; every finding traces to real tool output
Mission engine + War Room + Op Admiral ✅ Stable keyless through a connected local agent
Arsenal, MCP server, HTTP API ✅ Stable 36 built-in tools by default; 111 with the opt-in T3MP3ST_FULL_ARSENAL (+75 adapters, with dangerous/catalog-only drivers — metasploit, hydra, pacu, frida — behind narrow approved paths rather than generic execution) — both counts re-derive via verify-claims. security_recon over MCP
Egress-scope containment ✅ Stable (on by default) once a mission target is set, built-in networked tools refuse off-scope public hosts — not the target/subdomains, not loopback/private (SCOPE DENIED) — a tightened default, not a bare tool runner
Coordinated-disclosure pipeline ✅ Stable OSV novelty + live PoC + refuter panel + CVSS; drafts only, a human sends
White-box source analysis ⚠️ Experimental Multi-language ingest via web-tree-sitter (Python/JS/TS/Go/Java/C/C++); Python retains its regex parser, while other languages fail open to no extracted blocks; multi-model decomposition costs more tokens, not fewer
DeFi (Damn Vulnerable DeFi) ⚠️ Experimental reproduces known exploit classes; not novel discovery
Exploiter / Infiltrator / Exfiltrator / Ghost ⚠️ Experimental run the real tool-backed ReAct loop (same engine as recon); unproven as a coordinated swarm — single-agent is the benchmarked path, live swarm exploitation still unreliable
Advanced modules (cloud, persistence, swarm, cognition) 🚧 Planned interface-only in src/stubs/
Self-improvement loop 🧪 Research records lessons + proposals today; feeding them back into planning is roadmap

Full feature-by-feature breakdown: FEATURES.md.

Coverage by domain

Where the storm reaches today — and where it's headed. Same discipline as everything else: a domain is ✅ only when there's a receipt behind it.

Domain What it covers Status
🕸️ Web apps, APIs, auth flows, OWASP Top 10 ✅ Core — XBEN 90.1% pass@1
📂 Code white-box source audits, SAST-style vuln hunting ✅ Proven (hunt result) — held-out CVE-Zero: single-agent 8/10 exact file/line/CWE, 10/10 found (7 languages); the repo-ingest engine itself is still ⚠️ experimental
🚩 CTF wargames, practice ranges, challenges ✅ Proven — Cybench 23/40 hint-free
🔌 Network / Infra recon, service/stack fingerprinting; lateral + privesc ✅ recon (live nmap/DNS/HTTP engine) · ⚠️ lateral/privesc experimental
🤖 Embedded / IoT / OT firmware, robotics, ICS/SCADA OSS ✅ CVE pipeline live — coordinated-disclosure drafts held for vendors
📦 Supply chain dependency audits, install-without-confirmation ⚠️ Real — dedicated class; hit a CWE-829 on the held-out set
💰 Blockchain smart contracts, DeFi, Solidity ⚠️ Reproduction only — Damn Vulnerable DeFi, not novel discovery
☁️ Cloud AWS/GCP/Azure misconfig, IAM, serverless 🚧 In development
📱 Mobile Android/iOS app security 🚧 In development
🏢 Identity / AD Kerberos, pass-the-hash, AD attacks 🚧 In development
🔐 Binary / RE overflows, ROP, exploit dev 🚧 In development — needs specialized tooling

The class/squad architecture means new domains compose rather than fork — each is a loadout (specialist classes + arsenal + target adapter + a benchmark). 🚧 domains ship dark until they have a number.

Benchmarks

Headline results. Each recomputes from the committed JSON with npm run verify-claims; full methodology and caveats are in the linked docs.

Suite Result Context
XBEN — XBOW's 104-challenge suite, black-box pass@1 mean 90.1% (Wilson-95 86.2–92.9), floor 91/104 · gpt-5.5 XBOW self-reports 85% on the same suite; ours re-derives the graded verdict from committed artifacts (raw transcripts stripped for privacy)
XBEN — white-box (reported separately) pass@1 98.7%, best-ball 104/104 · gpt-5.5 never blended with the black-box number
Cybench — 40-task academic bench, Opus 4.8, no hints 23/40 (58%) hint-free, single-run pass@1 (verify-claims-enforced) not the raw-score record (Anthropic: 76.5% pass@10); every flag graded against the committed oracle
Cybench model matrix — identical 15-task committed subset, pass@1 Opus 4.7 vs 4.8, with per-task source receipts and separate failure/abstention/infrastructure outcomes rebuild and inspect the model/harness matrix; historical system comparison, not an isolated model ranking
CVE-Zero — 10 real post-cutoff (2026) CVEs, held-out, 7 languages single-agent 8/10 exact file/line/CWE (verified all-exact, stable) · 10/10 found (full pack) memorization- & fitting-proof: post-cutoff, and the hardened prompts were never tuned on these; verify-claims recomputes it. n=10, directional; the swarm's edge here is recall, not a coordination-beats-solo proof

How to read these:

  • Every solved flag is graded against a committed ground-truth oracle — not a self-report — and verify-claims recomputes the pass/fail. Raw per-step transcripts are stripped for operator privacy, so you re-check the graded verdict, not the raw tool output. Zero fabricated, enforced by an anti-fitting guard that runs on every push.
  • Black-box (source withheld) and white-box (source staged) are reported separately and never blended.
  • These ran a single-agent ReAct loop, not the 8-operator swarm. The swarm is framework architecture; it is not what scored these numbers.
  • Results are system-vs-system: this harness driving a strong current model, not an isolated-harness claim.

The number isn't the flex — the receipt is. A keyless, open-source harness that hands you the re-run instead of asking you to trust it: clone it, run npm run verify-claims, and every verdict above recomputes from its committed oracle in front of you.

Deeper reading: WALL_FORENSICS (per-challenge misses), CYBENCH, INTEGRITY_LEDGER (contamination audit and every retraction), OBSIDIVM (our own live web range).

Documentation

Doc Contents
Docs index operator, developer, benchmark, and release documentation map
Getting Started install, first launch, first safe mission, CLI basics, updates, and troubleshooting
Developer Guide source map, scripts, SDK usage, extension points, and release checks
API Reference local HTTP API route groups and integration notes
MCP Guide MCP server setup and security_recon usage
FEATURES.md feature-by-feature status ([x] shipped / [~] partial / [ ] planned)
SCOPE_AND_AUTHORIZATION authority model, scope receipts, evidence and retest rules
AUTHENTICATED_WORKFLOWS current authenticated-request support, capability diagnostics, and the design boundary for interactive signup
VERIFIED_PROVENANCE how findings become tool-proven instead of model-asserted
CONTRIBUTION_RECEIPTS PR receipt template for scope, run mode, model/harness labels, redaction, and verification
MODEL_MATRIX Reproducible cross-model benchmark matrix and arbitrary variant-test model selection
TEAM_PREVIEW first-run path and review script
INSTALL_MATRIX macOS / Linux readiness table
ARSENAL_ACTIVATION_PLAN optional external-tool setup
PULL_REQUEST_DELIVERY contributor and maintainer checklist for scoped, reviewable PRs
CYBENCH · WALL_FORENSICS · INTEGRITY_LEDGER · COGNITIVE_ARCHITECTURE benchmark methodology
RELEASE_CHECKLIST the gates a release must pass

Architecture

┌─────────────────────────────────────────────────────────────────┐
│                        T3MP3ST COMMAND                          │
├─────────────────────────────────────────────────────────────────┤
│   MISSION CONTROL  ◄──  TARGET MODEL  ──►  ARSENAL (TOOLS)       │
│                          ▲                                       │
│   AGENT CELL:  RECON · SCANNER · EXPLOITER · INFILTRATOR ·       │
│                EXFILTRATOR · GHOST · COORDINATOR · ANALYST       │
│                          ▲                                       │
│   EVIDENCE VAULT  ·  CREDENTIAL STORE  ·  FINDINGS LEDGER        │
│                          ▲                                       │
│   OPSEC LAYER  ·  COMMS CHANNEL  ·  LLM BACKBONE                 │
└─────────────────────────────────────────────────────────────────┘

Operators map to MITRE ATT&CK and Cyber Kill Chain phases (recon is live; later phases are scaffolded):

Operator Phase MITRE Function
Recon Reconnaissance TA0043 OSINT, network discovery, asset enumeration
Scanner Discovery TA0007 vulnerability scanning, service fingerprinting
Exploiter Initial Access TA0001 exploitation, payload delivery
Infiltrator Lateral Movement TA0008 post-exploitation, privilege escalation
Exfiltrator Collection / Exfil TA0009/10 data extraction, credential harvesting
Ghost Persistence TA0003 persistence, stealth, cleanup
Coordinator Command & Control TA0011 mission control, orchestration
Analyst Analysis — pattern analysis, reporting

Providers: OpenRouter, Venice, Anthropic, OpenAI, or a keyless local agent (Claude Code / Codex / Hermes / OpenCode / Oh My Pi). Set OPENROUTER_API_KEY / VENICE_API_KEY / ANTHROPIC_API_KEY, or connect an agent in Settings.

Integrations: node dist/mcp-server.js exposes security_recon to MCP-aware agents. npm run server starts the HTTP API (POST /api/mission/start, GET /api/mission/status, and more). Full reference in docs/.

Contributing — join the swarm

Red-teaming shouldn't be a priesthood. Bring an adapter, a prompt pack, a runbook, a new arsenal tool, or a bug report.

One rule, non-negotiable: everything here is for authorized testing only. Owned, scoped, or consenting targets. Build for defenders, or don't build it here.

  1. Fork it, branch it.
  2. Open a PR with tests. If you touch a headline number, npm run verify-claims has to stay green.

Release process and gates: RELEASE_CHECKLIST.

License

AGPL-3.0. See LICENSE.


Fortes fortuna iuvat — fortune favors the bold.

⊰•-•✧ LOVE PLINY ✧•-•⊱ 🌩️

View on GitHub

Recent activity

commits and pull requests

Recent open issues

view all

Releases and announcements

1 total
  1. v1.0.0 — Certified source checkpointv1.0.0Sep 8, 2026185 downloads

    This release marks a source commit that passed T3MP3ST's deterministic release quality gates. The single attached asset is the certified source ZIP. It was created by the tag workflow, extracted for a clean locked install and build, and verified against its SHA-256 and signed SLSA provenance before publication. Changes include reliable mission-control acknowledgements and status handling, configuration isolation for local verification, expanded release checks, and dependency updates with zero reported vulnerabilities. Certification covers lint (0 errors; 164 existing warnings permitted), type checking, 988 tests, configured coverage thresholds, deterministic contract suites, all four isolated local-server smoke suites, documentation checks, build, and package dry run. Optional live-provider, external Cybench corpus, PowerShell, and hardware checks are not certified by these results. Stop prevents new scheduling; requests and tools already in flight may finish. Commit: `29824d5625ede419ac8cdae418c8f4c72c6270f7`. Local certification note: unrestricted worker concurrency on a busy workstation produced two five-second test timeouts in separate tests. The complete local gate passe

Code frequency

additions and deletions
+156.3K-156.3KWeek of 2026-06-28: +156,286 linesWeek of 2026-06-28: -57,709 linesWeek of 2026-07-05: +4,594 linesWeek of 2026-07-05: -413 linesWeek of 2026-07-12: +6,333 linesWeek of 2026-07-12: -583 linesWeek of 2026-07-19: +13,964 linesWeek of 2026-07-19: -369 linesWeek of 2026-07-26: +2,463 linesWeek of 2026-07-26: -162 linesWeek of 2026-08-02: +4,756 linesWeek of 2026-08-02: -1,585 linesWeek of 2026-08-09: +73 linesWeek of 2026-08-09: -0 linesWeek of 2026-08-16: +0 linesWeek of 2026-08-16: -0 linesWeek of 2026-08-23: +1,356 linesWeek of 2026-08-23: -125 linesWeek of 2026-08-30: +9,765 linesWeek of 2026-08-30: -299 linesWeek of 2026-09-06: +849 linesWeek of 2026-09-06: -160 linesWeek of 2026-09-13: +0 linesWeek of 2026-09-13: -0 linesWeek of 2026-09-20: +0 linesWeek of 2026-09-20: -0 linesWeek of 2026-09-27: +0 linesWeek of 2026-09-27: -0 linesJun 28, 2026Sep 27, 2026
+200.4K lines added, -61.4K removed over the last year.

Commits per week

last 52 weeks
530Week of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 0 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 23 commitsWeek of 2026-07-05: 53 commitsWeek of 2026-07-12: 25 commitsWeek of 2026-07-19: 11 commitsWeek of 2026-07-26: 12 commitsWeek of 2026-08-02: 1 commitsWeek of 2026-08-09: 1 commitsWeek of 2026-08-16: 0 commitsWeek of 2026-08-23: 5 commitsWeek of 2026-08-30: 27 commitsWeek of 2026-09-06: 2 commitsWeek of 2026-09-13: 0 commitsWeek of 2026-09-20: 0 commitsWeek of 2026-09-27: 0 commitsOct 5, 2025Sep 27, 2026
160 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 1 commitsSun 10:00 — 0 commitsSun 11:00 — 0 commitsSun 12:00 — 4 commitsSun 13:00 — 0 commitsSun 14:00 — 3 commitsSun 15:00 — 0 commitsSun 16:00 — 1 commitsSun 17:00 — 6 commitsSun 18:00 — 1 commitsSun 19:00 — 0 commitsSun 20:00 — 2 commitsSun 21:00 — 9 commitsSun 22:00 — 0 commitsSun 23:00 — 1 commitsMon 0:00 — 2 commitsMon 1:00 — 3 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 1 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 1 commitsMon 9:00 — 0 commitsMon 10:00 — 4 commitsMon 11:00 — 0 commitsMon 12:00 — 1 commitsMon 13:00 — 0 commitsMon 14:00 — 2 commitsMon 15:00 — 0 commitsMon 16:00 — 0 commitsMon 17:00 — 4 commitsMon 18:00 — 5 commitsMon 19:00 — 2 commitsMon 20:00 — 0 commitsMon 21:00 — 2 commitsMon 22:00 — 0 commitsMon 23:00 — 1 commitsTue 0:00 — 1 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 0 commitsTue 10:00 — 0 commitsTue 11:00 — 1 commitsTue 12:00 — 2 commitsTue 13:00 — 1 commitsTue 14:00 — 0 commitsTue 15:00 — 0 commitsTue 16:00 — 0 commitsTue 17:00 — 0 commitsTue 18:00 — 0 commitsTue 19:00 — 0 commitsTue 20:00 — 0 commitsTue 21:00 — 0 commitsTue 22:00 — 0 commitsTue 23:00 — 0 commitsWed 0:00 — 0 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 0 commitsWed 9:00 — 0 commitsWed 10:00 — 1 commitsWed 11:00 — 1 commitsWed 12:00 — 0 commitsWed 13:00 — 1 commitsWed 14:00 — 0 commitsWed 15:00 — 0 commitsWed 16:00 — 1 commitsWed 17:00 — 3 commitsWed 18:00 — 4 commitsWed 19:00 — 6 commitsWed 20:00 — 4 commitsWed 21:00 — 0 commitsWed 22:00 — 0 commitsWed 23:00 — 1 commitsThu 0:00 — 1 commitsThu 1:00 — 2 commitsThu 2:00 — 1 commitsThu 3:00 — 0 commitsThu 4:00 — 4 commitsThu 5:00 — 3 commitsThu 6:00 — 0 commitsThu 7:00 — 1 commitsThu 8:00 — 0 commitsThu 9:00 — 0 commitsThu 10:00 — 1 commitsThu 11:00 — 1 commitsThu 12:00 — 1 commitsThu 13:00 — 6 commitsThu 14:00 — 12 commitsThu 15:00 — 0 commitsThu 16:00 — 5 commitsThu 17:00 — 0 commitsThu 18:00 — 0 commitsThu 19:00 — 1 commitsThu 20:00 — 0 commitsThu 21:00 — 1 commitsThu 22:00 — 0 commitsThu 23:00 — 2 commitsFri 0:00 — 1 commitsFri 1:00 — 2 commitsFri 2:00 — 1 commitsFri 3:00 — 1 commitsFri 4:00 — 0 commitsFri 5:00 — 3 commitsFri 6:00 — 0 commitsFri 7:00 — 1 commitsFri 8:00 — 0 commitsFri 9:00 — 0 commitsFri 10:00 — 1 commitsFri 11:00 — 5 commitsFri 12:00 — 0 commitsFri 13:00 — 0 commitsFri 14:00 — 0 commitsFri 15:00 — 0 commitsFri 16:00 — 2 commitsFri 17:00 — 2 commitsFri 18:00 — 1 commitsFri 19:00 — 0 commitsFri 20:00 — 1 commitsFri 21:00 — 1 commitsFri 22:00 — 2 commitsFri 23:00 — 1 commitsSat 0:00 — 0 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 2 commitsSat 6:00 — 1 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 1 commitsSat 12:00 — 0 commitsSat 13:00 — 0 commitsSat 14:00 — 0 commitsSat 15:00 — 0 commitsSat 16:00 — 1 commitsSat 17:00 — 0 commitsSat 18:00 — 0 commitsSat 19:00 — 1 commitsSat 20:00 — 1 commitsSat 21:00 — 1 commitsSat 22:00 — 0 commitsSat 23:00 — 2 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits40 (20%)
Community commits164 (80%)

204 commits in total over the last year.

DateListRankStars gained
Jul 5, 2026daily#4+37
  • freeCodeCamp/freeCodeCamp

    freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming, and computer science for free.

    456.7K stars · TypeScript

  • openclaw/openclaw

    The AI that really does things. Any OS. Any Platform. The lobster way. 🦞

    391.3K stars · TypeScript

  • obra/superpowers

    An agentic skills framework & software development methodology that works.

    295.2K stars · Shell

  • NousResearch/hermes-agent

    The agent that grows with you

    251.2K stars · Python

  • anomalyco/opencode

    The open source coding agent.

    211.7K stars · TypeScript

  • n8n-io/n8n

    Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

    206.7K stars · TypeScript