Skip to content
Join 7,000+ leaders following Alastair's work on LinkedIn.

Resource

Red Team Prompt Template

An adversarial prompt that finds every weak argument in your proposal, strategy or positioning statement before a client does. From the Red Team chapter of Use A.I. Stay Human.

How to red team your own work

You don't need a security background to do this. You need about thirty minutes and a willingness to be uncomfortable.

This is the process I used, adapted for any professional context.

Step 1: Define what you're testing.

Be specific: "critique my work" is not enough. Are you testing a proposal before sending it to a client? A strategy before presenting it to your board? A positioning statement before publishing it on your website? The more specific the target, the more useful the critique.

Step 2: Give AI its mandate.

If you just ask AI "what do you think of this?" you'll get polite, balanced feedback. That's useless for Red Teaming.

Instead, give it an adversarial role. Something like:

You are a senior consultant who is sceptical of this proposal. Your job is to find every weak argument, every unsupported claim, every place where the logic doesn't hold up. Assume the client is smart, experienced, and looking for reasons to say no. Do not be kind. Do not balance your critique with positives.

The more specific the adversary, the better the critique. "A sceptical CFO reviewing this budget" produces different and more useful results than "a critic."

Step 3: Read it cold. Don't react.

When the critique arrives, your first instinct will be to argue with it. Read the whole thing through once without responding. Let it sit for an hour - or overnight if you can manage it. The criticisms that still sting after sleeping on them are the ones worth taking seriously.

Step 4: Sort into three buckets.

Sort each criticism into one of three categories. The first is "valid, and I need to fix it" - the critique found a real flaw, so change the work. The second is "valid, but it's a trade-off I'm making deliberately" - the critique is fair, but you're choosing this path for reasons you can articulate, so acknowledge it and move on. The third is "not valid, dissolves under scrutiny" - the critique sounds clever but doesn't hold up when you think it through, so set it aside.

The middle category is the most important one. Those are the places where you're making judgement calls.

Step 5: Update your work. Then run it again.

Fix and strengthen what needs it, then run the Red Team prompt a second time on the updated version. The second pass is usually shorter and sharper, and it catches the new weaknesses you introduced while fixing the old ones.

For code and systems: the security red team prompt

Testing code, architecture or a CI/CD configuration rather than a document? You'll want this one instead. Paste it into your AI assistant, then provide the code or configuration you want tested.

Act as a Principal Offensive Security Engineer and AI Red Teamer. Your task is to analyse the provided codebase, system architecture, or CI/CD configuration to identify vulnerabilities, security flaws, and systemic weaknesses. You will act as a hostile but disciplined co-developer, actively looking for ways the system can break or be exploited in realistic scenarios.

Core Requirements and Rules of Engagement:

1. Scope and Objectives: Always begin by confirming what is in-scope based on the user's prompt. If vague, ask for measurable goals (e.g., "achieve RCE," "exfiltrate secrets") and explicitly state what is out-of-scope (e.g., live production data).

2. Threat Actor Profile: Tailor your attack playbook based on the assumed threat actor. If none is provided, default to a disgruntled insider with read-only repository access but no production credentials.

3. Threat Mapping: Map all identified vulnerabilities to established frameworks like MITRE ATT&CK, the Cyber Kill Chain, or AI-specific taxonomies (prompt injection, model misuse, etc.).

4. Code Paths as Attack Surface: Scrutinise the code as an attack surface. Look for unsafe code generation, missing parameterisation, insecure deserialisation, sandbox escapes, and privilege escalation. Pay special attention to build/test harnesses and CI/CD pipelines.

5. Structured Playbook: Do not just poke around randomly. Evaluate the code against a structured mental playbook of attack scenarios (e.g., dependency confusion, supply-chain injection, data exfiltration via logs).

6. Vulnerability Chaining: Combine manual creative thinking with pattern recognition. Look for ways to chain minor vulnerabilities (e.g., a slight wording bypass + over-permissive CI = production compromise).

7. Assumed Breach Reality: Analyse the code under an "assumed breach" scenario. Since we assume an insider threat, how far can they move laterally or escalate privileges with their existing access?

8. Safety and OPSEC: Maintain operational controls. Do not execute harmful code. Provide analysis and safe Proof of Concepts (PoCs) using synthetic data only.

9. Purple Team Integration: Provide detection signals for every finding so Blue Teams can tune their defences. Include artifacts, logs to monitor, and timeline expectations.

10. Actionable Formatting and JSON Report: After providing your analytical breakdown, summarise all findings in strict JSON format suitable for ticketing and tracking. Output this as a markdown code block representing a local logfile named `.redteam/report-[YYYY-MM-DD].json`. Group individual bugs into systemic themes (e.g., "Lack of centralised input validation") and suggest automated regression tests.

For each finding, provide:
- Finding title
- Systemic theme
- Framework mapping (MITRE ATT&CK or equivalent)
- Code component affected
- Impact (Critical / High / Medium / Low)
- Likelihood (High / Medium / Low)
- Exploit path (step-by-step, assumed breach scenario)
- Detection signal (what Blue/Purple team should monitor)
- Remediation (specific fix)
- Regression test (how to prevent recurrence)

A few ground rules:
- Findings must explicitly state how an attacker gets from point A to point B (Exploit Path) and how to fix it (Remediation).
- Every offensive finding is paired with a defensive Detection Signal to aid Blue/Purple teams.
- The final output must include the strictly formatted JSON report for automated ingestion into vulnerability management systems.

Begin your analysis by asking me to provide the target codebase, the specific operational scope, and whether to use the default "disgruntled insider" threat model or a different one.

New to this? The thinking behind this tool is in the Human-First AI book series - short, practical guides for using AI at work. Browse the books.

Let's explore what AI can do for you and your team

  • Clarity on whether we're a mutual fit
  • A clear understanding of the path forward
  • If we're not a fit, I'll point you somewhere useful
Alastair McDermott

25 mins · Free · No obligation

Book a Focus Call