CB security-analyst
Use when the user wants a security audit, penetration test, threat model, vulnerability hunt, security fix plan, SBOM, compliance mapping, privacy assessment, or security posture comparison between runs.
As a process B 79/100 · Nearly there — no weak spots found
How to improve
- The SKILL.md body is over 5,000 tokens: move reference detail into references/ and load it when needed.
For the model run — optional
- Your own cases (evals/evals.json, 4–6 real requests with expected answers): the full check would then run those instead of a model-drafted suite.
- A spec.yaml with trigger phrases and assertions — a behaviour contract for CI; `skilltest init` writes a template.
Guard findings · 40
✓ No critical or high findings
Medium and low: 40
-
low Risky intent
intent-offensive-securityassets/templates/final-report.md:11Offensive-security / dual-use content (legitimate for authorised testing; review intended use)**Methodology:** Offensive penetration testing — multi-phase analysis with team-based orchestration
-
low Risky intent
intent-offensive-securityassets/templates/final-report.md:86Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (quoted — discussed, not commanded)Map all findings to MITRE ATT&CK Enterprise tactics to visualize kill chain exposure and identify blind spots. Only include tactics relevant to the project's architecture (e.g., skip Lateral Movement
quoted -
low Risky intent
intent-offensive-securityassets/templates/final-report.md:93Offensive-security / dual-use content (legitimate for authorised testing; review intended use)| Privilege Escalation | TA0004 | | |
-
low Risky intent
intent-offensive-securityassets/templates/finding.md:78Offensive-security / dual-use content (legitimate for authorised testing; review intended use)## Exploit Chain
-
low Risky intent
intent-offensive-securityREADME.md:103Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (quoted — discussed, not commanded)- Use phrases like "run a security audit", "penetration test", "threat model", "SBOM", "compliance mapping"
quoted -
low Risky intent
intent-offensive-securityreferences/commands/full.md:2Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (quoted — discussed, not commanded)description: "Complete 9-phase offensive security analysis. No prompts — runs everything."
quoted -
low Risky intent
intent-offensive-securityreferences/commands/security-analysis.md:42Offensive-security / dual-use content (legitimate for authorised testing; review intended use)| **Elevation of Privilege** | Can access controls at this boundary be bypassed? (privilege escalation, role confusion, token scope abuse) |
-
low Risky intent
intent-offensive-securityreferences/commands/security-analyst.md:2Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (quoted — discussed, not commanded)description: "Deep offensive security analysis — penetration tester mindset across 9 phases with agent-based orchestration. Finds real vulnerabilities, not checkbox compliance."
quoted -
low Risky intent
intent-offensive-securityreferences/commands/security-analyst.md:6Offensive-security / dual-use content (legitimate for authorised testing; review intended use)# Security Analyst — Offensive Penetration Testing Orchestrator
-
low Risky intent
intent-offensive-securityreferences/commands/security-analyst.md:8Offensive-security / dual-use content (legitimate for authorised testing; review intended use)You are the orchestrator for an agent-based offensive security analysis. You coordinate specialized agents through multi-stage execution, consolidating findings and managing the analysis lifecycle.
-
low Risky intent
intent-offensive-securityreferences/commands/security-analyst.md:12Offensive-security / dual-use content (legitimate for authorised testing; review intended use)1. **Offensive framing always.** Every analysis starts with "If I were attacking this system..." — never "Verify that...". You are leading a red team.
-
low Risky intent
intent-offensive-securityreferences/commands/security-analyst.md:16Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (detector / deny-list definition)5. **Business logic over syntax.** The most dangerous bugs are logic flaws: race conditions, decision manipulation, privilege escalation, action hijacking.
detector -
low Risky intent
intent-offensive-securityreferences/commands/threat-model.md:60Offensive-security / dual-use content (legitimate for authorised testing; review intended use)- Planning a full penetration test
-
low Risky intent
intent-offensive-securityreferences/plan-lod.md:468Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (quoted — discussed, not commanded)description: "Complete 9-phase offensive security analysis. No prompts — runs everything."
quoted -
low Risky intent
intent-offensive-securityreferences/plugins/aws-gcp-azure.md:93Offensive-security / dual-use content (legitimate for authorised testing; review intended use)2. Are there cross-service permissions that allow lateral movement? (Lambda with S3 full access)
-
low Risky intent
intent-offensive-securityreferences/plugins/python-django.md:68Offensive-security / dual-use content (legitimate for authorised testing; review intended use)- Can users be assigned to groups via API endpoints? (privilege escalation if an endpoint allows group modification)
-
low Risky intent
intent-offensive-securityreferences/plugins/README.md:57Offensive-security / dual-use content (legitimate for authorised testing; review intended use)| `logic-authz-escalation` | Privilege Escalation | Framework-specific privilege escalation vectors |
-
low Risky intent
intent-offensive-securityreferences/prompts/container-security.md:3Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (quoted — discussed, not commanded)You are a penetration tester auditing the project's containerization and orchestration configuration for escape vectors, privilege escalation, and runtime security gaps.
quoted -
low Risky intent
intent-offensive-securityreferences/prompts/container-security.md:51Offensive-security / dual-use content (legitimate for authorised testing; review intended use)1. **Privilege Escalation:**
-
low Risky intent
intent-offensive-securityreferences/prompts/container-security.md:127Offensive-security / dual-use content (legitimate for authorised testing; review intended use)- Privilege escalation findings should specify the exact capability or permission that enables it
-
low Risky intent
intent-offensive-securityreferences/prompts/exploit-developer.md:68Offensive-security / dual-use content (legitimate for authorised testing; review intended use)- IDOR + state manipulation → privilege escalation
-
low Risky intent
intent-offensive-securityreferences/prompts/finding-critic.md:3Offensive-security / dual-use content (legitimate for authorised testing; review intended use)You are a skeptical senior security engineer reviewing findings from a penetration testing team. Your job is to challenge every finding, catch false positives, validate exploitability claims, verify f
-
low Risky intent
intent-offensive-securityreferences/prompts/finding-critic.md:77Offensive-security / dual-use content (legitimate for authorised testing; review intended use)For each exploit chain:
-
low Risky intent
intent-offensive-securityreferences/prompts/logic-authz-escalation.md:1Offensive-security / dual-use content (legitimate for authorised testing; review intended use)# Business Logic Agent — Authorization & Privilege Escalation
-
low Risky intent
intent-offensive-securityreferences/prompts/logic-authz-escalation.md:3Offensive-security / dual-use content (legitimate for authorised testing; review intended use)You are a penetration tester tracing every authorization decision for bypass opportunities, privilege escalation, and cross-user access.
-
low Risky intent
intent-offensive-securityreferences/prompts/recon-agent.md:3Offensive-security / dual-use content (legitimate for authorised testing; review intended use)You are a reconnaissance scout for an offensive security analysis. Your job is to rapidly map one aspect of the target codebase and produce a structured LOD-2 section file.
-
low Risky intent
intent-offensive-securityreferences/prompts/recon-agent.md:143Offensive-security / dual-use content (legitimate for authorised testing; review intended use)- Glob for security docs: **/security*.md, **/threat-model*, **/pentest*
-
low Risky intent
intent-offensive-securityreferences/prompts/threat-model-agent.md:3Offensive-security / dual-use content (legitimate for authorised testing; review intended use)You are a threat modeling specialist on an offensive security team. Your job is to take the reconnaissance report and produce a structured threat model that maps every trust boundary to concrete threa
-
low Risky intent
intent-offensive-securityreferences/prompts/threat-model-agent.md:36Offensive-security / dual-use content (legitimate for authorised testing; review intended use)| **Elevation of Privilege** | Can access controls at this boundary be bypassed? (privilege escalation, role confusion, token scope abuse) |
-
low Risky intent
intent-offensive-securityreferences/prompts/threat-model-agent.md:92Offensive-security / dual-use content (legitimate for authorised testing; review intended use)1. List which ATT&CK tactics (Initial Access, Execution, Persistence, Privilege Escalation, Defense Evasion, Credential Access, Discovery, Collection, Exfiltration, Impact) are relevant given the arch
-
low Risky intent
intent-offensive-securityreferences/prompts/threat-model-agent.md:95Offensive-security / dual-use content (legitimate for authorised testing; review intended use)4. Skip tactics that don't apply to the project's architecture (e.g., Lateral Movement for a single-service app, Persistence for a stateless function)
-
low Risky intent
intent-offensive-securitySKILL.md:3Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (quoted — discussed, not commanded)description: "Use when the user wants a security audit, penetration test, threat model, vulnerability hunt, security fix plan, SBOM, compliance mapping, privacy assessment, or security posture compari
quoted -
low Risky intent
intent-offensive-securitySKILL.md:16Offensive-security / dual-use content (legitimate for authorised testing; review intended use)- When a security audit or penetration test is requested
-
low Risky intent
intent-offensive-securitySKILL.md:309Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (detector / deny-list definition)- **Absolute file paths with line numbers**: Recon agents record exact file locations so downstream agents can `Read` specific code without re-scanning. This is an efficiency optimization, not a privi
detector -
low Risky intent
intent-offensive-securitySKILL.md:344Offensive-security / dual-use content (legitimate for authorised testing; review intended use)2. **Restrict access** — Treat run directories with the same access controls as penetration test reports. Share on a need-to-know basis.
-
low Risky intent
intent-offensive-securitySKILL.md:350Offensive-security / dual-use content (legitimate for authorised testing; review intended use)The skill sets `always: false` in its frontmatter. It only activates when the user explicitly requests a security audit, penetration test, threat model, vulnerability scan, SBOM, compliance mapping, o
A further 4 matches are quotations in this security skill's documentation and are not counted as findings.
Files scanned: 74. Evidence is masked. Grey chips explain why severity was lowered.
Against the Agent Skills spec
- warning
body-longSKILL.md body ≈ 5340 tokens (recommended < 5000); move details to references/ - note
frontmatter-keyunknown frontmatter key "always"
Process rating: all ten parameters 79/100
- 60Tools and files. Uses tools (bash, web, python, node) that frontmatter does not declare
- 60Result and completion. Output format stated, no completion criterion
- 70When it triggers. States when to use, but not when not to
- 70Inputs and preconditions. Inputs and preconditions are listed
- 70Execution cost. Instruction body is 5340 tokens
- 100Steps. 35 steps
- 100Failures and branches. 6 branches, has a failure section
- 100Consistency. Name and required fields are in place
- 100Running it twice. Mutating operations check current state
- 100Progress reporting. Reports progress
- medium Safety rules and hard prohibitions inside a skill: they belong in the system prompt, here they protect nothing
- low 10 top-level sections: this looks like several domains in one skill
Everything here is measured from the skill text rather than judged by a model, so the numbers are checkable. A parameter weighs more when it is a more common reason for the process to stall.
Quality signals
- +5Description has no quoted example phrases that should trigger the skill
- +4Description does not say when NOT to use the skill (false activations)
- +1No license
- +2Single-language instructions
- +3Description length 203: enough signal without eating the budget
- +4Structure: 35 headings
- +3Step-by-step instructions: 35 items
- +3Output format is stated explicitly
- +4Has examples (12 code blocks)
- +4Reference files are cited in the instructions (3 of 5)
Quality base 70; lint remarks subtract, signals add up to 100. Result: 80.
External checks
ClawHub: suspicious
This appears to be a legitimate but very powerful security-audit skill, with review concerns because it can read secrets, generate exploit details, and in one plugin reaches beyond the stated project-only scope.
LLM: suspicious (high) · VirusTotal: suspicious · 28 May 2026