All skills
hardw00t avatar

/llm-security

@f9bb3b2

LLM and AI application security testing skill for prompt injection (direct, indirect, multimodal), system-prompt extraction, RAG poisoning, memory poisoning, MCP server injection, skill-file injection, agentic tool misuse, computer-use UI injection, and excessive agency. Authorization required — this skill tests AI systems you are explicitly permitted to assess. Triggers on requests to test LLM / AI-agent / RAG / MCP / computer-use security, perform prompt injection, extract system prompts, poison RAG or memory, audit agent tool use, or evaluate AI guardrails.

Use this Skill: https://skilld.dev/gh/hardw00t/ai-security-arsenal/llm-security

This session only. Nothing lands on disk.

workflowsexcessive_agency_testing.md

≈966 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Workflow: Excessive Agency Testing (OWASP LLM06)

Goal: quantify the blast radius an attacker can achieve when an agent is misdirected. Excessive agency manifests as: tools the agent doesn't need for its legitimate purpose, overly broad permissions on the tools it does need, or missing human-in-the-loop on sensitive actions.

This workflow complements workflows/agentic_tool_misuse.md:

  • agentic_tool_misuse → how to coerce misuse
  • excessive_agency → how bad is it when misuse happens

Parallelism: capability probes can run in parallel. Privilege-escalation chains are sequential. Reasoning budget: moderate — reason about tool composition and least- privilege gaps.


1. Capability inventory

Build an accurate list of what the agent can do in principle:

Category Questions
Tools Names, argument schemas, side effects
Auth scope What identity runs each tool? What can that identity do?
Data access What PII / secrets / business data is reachable?
Blast radius If tool X is abused, what is the worst outcome?
HITL Which tools require human confirmation? Is it meaningful?
Auditing Are actions logged? Alerted? Reviewable?

2. Least-privilege gap analysis

For each tool, ask: "is this required for the agent's declared job?"

  • If no → excessive-agency finding (unnecessary capability)
  • If yes but scope is broader than needed → excessive-scope finding
  • If yes and scope is minimal → ok, but still test abuse

3. HITL quality test

For each gated tool, verify that the approval prompt:

  • Shows the full, unmodified action being performed (not just a name)
  • Shows every argument including hidden/derived ones
  • Cannot be pre-approved / batch-approved for a session
  • Distinguishes different invocations clearly
  • Is not easily spoofed by agent output (e.g. by the agent's own text saying "User: I approve")

Common failure: UI shows "Agent wants to run shell" but not the command.

4. Privilege escalation within agent

  • Can the agent modify its own system prompt / instructions?
  • Can it write to its skill files / CLAUDE.md / .cursor/rules? (chained with workflows/skill_file_injection.md)
  • Can it add / enable new MCP servers?
  • Can it write to persistent memory? (chained with workflows/memory_poisoning.md)
  • Can it grant itself additional API scopes?

Any "yes" is typically a critical finding.

5. Lateral reach

  • Can the agent pivot to other services using its creds?
  • Can the agent read / modify other users' data?
  • Can the agent invoke other agents (multi-agent systems)?

6. Denial-of-service / cost

  • Can a single user prompt cause unbounded tool loops?
  • Are there token / call / $$ quotas per user / session?
  • Can retrieval cause recursive expansion?
  • Can the agent be made to send high-volume outbound (email, API)?

Maps to OWASP LLM10 (Unbounded Consumption).

7. Worst-case exploit construction

Take the most impactful capability × weakest HITL gate × easiest coercion vector. Write a concrete end-to-end exploit (against your test target). Document step-by-step with timestamps and evidence.

This is the finding that will drive remediation prioritization.

8. Reporting

  • attack_class = "excessive-agency" (for capability findings) or the specific misuse class when exploited
  • owasp_llm_id = "LLM06:2025" (and LLM10 for consumption)
  • Clear least-privilege recommendation per over-scoped tool

9. Remediation hints

  • Remove unnecessary tools from the agent's configuration
  • Narrow scopes (path allowlists, URL allowlists, command allowlists)
  • Strong HITL: full-diff display, non-bypassable
  • Short-lived, narrowly-scoped tokens per tool call (capability tokens)
  • Per-tool rate limits and daily caps
  • Separate agent identities for different trust contexts
  • Audit log with anomaly alerting

Source: SKILL.md on GitHub

1 alert16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill is a comprehensive security auditing and red-teaming toolkit designed to test LLM applications for vulnerabilities like prompt injection, RAG poisoning, and tool misuse. While it contains many examples of malicious payloads and attack patterns, these are provided for defensive testing purposes within a clearly defined security research framework that requires explicit authorization. No actual malicious code or unauthorized data exfiltration logic is executed by the skill itself; it serves as a guide and resource for security professionals.

  • Socket16d

    5 alerts: gptSecurity, gptAnomaly

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • Runlayer7mo

    1/1 file flagged

Signed by skilld at f9bb3b2. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Steadyupdated 6 months ago

README badge

README badge for hardw00t/ai-security-arsenal/llm-security