All skills
semgrep avatar

/llm-security

@327da93 official
by semgrepsemgrep/skills317 stars
31

Security guidelines for LLM applications based on OWASP Top 10 for LLM 2025. Use when building LLM apps, reviewing AI security, implementing RAG systems, or asking about LLM vulnerabilities like 'prompt injection' or 'check LLM security'. IMPORTANT: Always consult this skill when building chatbots, AI agents, RAG pipelines, tool-using LLMs, agentic systems, or any application that calls an LLM API (OpenAI, Anthropic, Gemini, etc.) — even if the user doesn't explicitly mention security. Also use when users import 'openai', 'anthropic', 'langchain', 'llamaindex', or similar LLM libraries.

Use this Skill: https://skilld.dev/gh/semgrep/skills/llm-security

This session only. Nothing lands on disk.

rulessupply-chain.md

≈2.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

LLM03: Secure LLM Supply Chain

LLM supply chains include pre-trained models, fine-tuning data, embeddings, plugins, and deployment infrastructure. Vulnerabilities can arise from compromised model repositories, malicious training data, vulnerable dependencies, or tampered model files.

Risk factors: Unverified model sources, malicious pickle files, compromised LoRA adapters, outdated dependencies, unclear licensing.


Model Verification

Vulnerable (unverified model download):

from transformers import AutoModel

# Downloading without verification
model = AutoModel.from_pretrained("random-user/suspicious-model")

Secure (verified model with integrity checks):

from transformers import AutoModel
import hashlib
import requests

TRUSTED_MODELS = {
    "meta-llama/Llama-2-7b-hf": {
        "sha256": "abc123...",  # Known good hash
        "license": "llama2",
        "verified_date": "2024-01-15"
    }
}

def verify_model_integrity(model_name: str, model_path: str) -> bool:
    """Verify model file integrity against known hashes."""
    if model_name not in TRUSTED_MODELS:
        raise ValueError(f"Model {model_name} not in trusted list")

    expected_hash = TRUSTED_MODELS[model_name]["sha256"]

    # Calculate hash of downloaded model
    sha256_hash = hashlib.sha256()
    with open(model_path, "rb") as f:
        for chunk in iter(lambda: f.read(4096), b""):
            sha256_hash.update(chunk)

    actual_hash = sha256_hash.hexdigest()
    return actual_hash == expected_hash

def load_verified_model(model_name: str):
    """Load model only from trusted sources with verification."""

    # Only allow models from trusted organizations
    trusted_orgs = ["meta-llama", "openai", "anthropic", "google", "microsoft"]
    org = model_name.split("/")[0] if "/" in model_name else None

    if org not in trusted_orgs:
        raise ValueError(f"Model organization {org} not trusted")

    # Use safe serialization (avoid pickle)
    model = AutoModel.from_pretrained(
        model_name,
        trust_remote_code=False,  # Never trust remote code
        use_safetensors=True,     # Use safe tensor format
    )

    return model

Safe Model Loading (Avoid Pickle Exploits)

Vulnerable (unsafe pickle loading):

import pickle
import torch

# DANGEROUS: Pickle can execute arbitrary code
with open("model.pkl", "rb") as f:
    model = pickle.load(f)

# Also dangerous
model = torch.load("model.pt")  # Uses pickle internally

Secure (safe tensor loading):

from safetensors import safe_open
from safetensors.torch import load_file
import torch

def load_model_safely(model_path: str):
    """Load model using safetensors format (no code execution)."""

    if model_path.endswith(".safetensors"):
        # Safetensors is safe - no arbitrary code execution
        tensors = load_file(model_path)
        return tensors

    elif model_path.endswith((".pt", ".pth", ".pkl", ".pickle")):
        # Pickle-based formats are dangerous
        raise ValueError(
            "Pickle-based model files (.pt, .pkl) can execute arbitrary code. "
            "Convert to safetensors format first."
        )

    else:
        raise ValueError(f"Unknown model format: {model_path}")

# For PyTorch models, use weights_only=True (Python 3.10+)
def load_pytorch_safely(model_path: str):
    """Load PyTorch model with restricted unpickler."""
    return torch.load(model_path, weights_only=True)

Dependency Management

Vulnerable (unpinned dependencies):

# requirements.txt
transformers
torch
langchain

Secure (pinned with hashes):

# requirements.txt - pinned versions with hashes
transformers==4.36.0 \
    --hash=sha256:abc123...
torch==2.1.0 \
    --hash=sha256:def456...
langchain==0.1.0 \
    --hash=sha256:ghi789...
# Use pip-audit to check for vulnerabilities
# pip-audit --requirement requirements.txt

# Generate SBOM for AI components
# cyclonedx-py requirements requirements.txt -o sbom.json

ML Bill of Materials (ML-BOM)

Implementation:

import json
from datetime import datetime

def generate_ml_bom(model_config: dict) -> dict:
    """Generate ML Bill of Materials for model tracking."""

    ml_bom = {
        "bomFormat": "CycloneDX",
        "specVersion": "1.5",
        "version": 1,
        "metadata": {
            "timestamp": datetime.utcnow().isoformat(),
            "component": {
                "type": "machine-learning-model",
                "name": model_config["name"],
                "version": model_config["version"]
            }
        },
        "components": [
            {
                "type": "machine-learning-model",
                "name": model_config["base_model"],
                "version": model_config["base_model_version"],
                "purl": f"pkg:huggingface/{model_config['base_model']}",
                "properties": [
                    {"name": "ml:model_type", "value": "llm"},
                    {"name": "ml:training_date", "value": model_config["training_date"]},
                    {"name": "ml:license", "value": model_config["license"]}
                ]
            }
        ],
        "dependencies": model_config.get("dependencies", []),
        "externalReferences": [
            {
                "type": "documentation",
                "url": model_config.get("model_card_url")
            }
        ]
    }

    return ml_bom

# Example usage
model_config = {
    "name": "my-fine-tuned-llm",
    "version": "1.0.0",
    "base_model": "meta-llama/Llama-2-7b-hf",
    "base_model_version": "2.0",
    "training_date": "2024-01-15",
    "license": "llama2",
    "model_card_url": "https://example.com/model-card"
}

bom = generate_ml_bom(model_config)

LoRA Adapter Security

Vulnerable (unverified adapter):

from peft import PeftModel

# Loading untrusted adapter
model = PeftModel.from_pretrained(base_model, "random-user/lora-adapter")

Secure (verified adapter loading):

from peft import PeftModel
import hashlib

TRUSTED_ADAPTERS = {
    "verified-org/safe-adapter": {
        "sha256": "abc123...",
        "base_model": "meta-llama/Llama-2-7b-hf",
        "verified_by": "security-team",
        "verified_date": "2024-01-15"
    }
}

def load_verified_adapter(base_model, adapter_name: str):
    """Load LoRA adapter only from trusted sources."""

    if adapter_name not in TRUSTED_ADAPTERS:
        raise ValueError(f"Adapter {adapter_name} not in trusted list")

    adapter_info = TRUSTED_ADAPTERS[adapter_name]

    # Verify adapter is compatible with base model
    if adapter_info["base_model"] != base_model.config._name_or_path:
        raise ValueError("Adapter not compatible with base model")

    # Load with safetensors
    model = PeftModel.from_pretrained(
        base_model,
        adapter_name,
        use_safetensors=True
    )

    return model

Vendor and Data Source Vetting

Implementation:

from dataclasses import dataclass
from enum import Enum
from typing import Optional
from datetime import datetime

class TrustLevel(Enum):
    VERIFIED = "verified"
    TRUSTED = "trusted"
    UNTRUSTED = "untrusted"

@dataclass
class DataSourceConfig:
    name: str
    url: str
    trust_level: TrustLevel
    license: str
    last_audit: datetime
    data_processing_agreement: bool

def validate_data_source(source: DataSourceConfig) -> bool:
    """Validate data source meets security requirements."""

    # Check trust level
    if source.trust_level == TrustLevel.UNTRUSTED:
        return False

    # Ensure recent security audit
    days_since_audit = (datetime.now() - source.last_audit).days
    if days_since_audit > 90:
        return False

    # Require DPA for training data
    if not source.data_processing_agreement:
        return False

    # Verify acceptable license
    acceptable_licenses = ["MIT", "Apache-2.0", "CC-BY-4.0", "public-domain"]
    if source.license not in acceptable_licenses:
        return False

    return True

Key Prevention Rules

  1. Verify model sources - Only use models from trusted organizations
  2. Use safe serialization - Prefer safetensors over pickle formats
  3. Pin dependencies - Use exact versions with hash verification
  4. Maintain ML-BOM - Track all model components and data sources
  5. Audit regularly - Review models and dependencies for vulnerabilities
  6. Verify adapters - Treat LoRA/PEFT adapters with same scrutiny as models
  7. Check licenses - Ensure compliance with all model and data licenses
  8. Never trust remote code - Set trust_remote_code=False

References:

Source: SKILL.md on GitHub

1 alert16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The skill provides comprehensive security guidelines and code examples for building secure LLM applications, based on the OWASP Top 10 for LLMs 2025. It serves as an educational resource to help developers mitigate risks like prompt injection and sensitive data exposure. While the skill contains examples of vulnerable code with hardcoded credentials, these are clearly labeled as insecure patterns for demonstration and educational purposes, using non-functional example values.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    12/14 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 327da93. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Steadyupdated 7 months ago
  • llm-security
  • prompt-injection
  • owasp
  • rag
  • ai-agents
  • data-poisoning
  • sensitive-disclosure
  • output-handling
  • vector-embeddings

README badge

README badge for semgrep/skills/llm-security

Provides security guidelines for LLM applications based on OWASP Top 10 for LLM 2025, covering prompt injection, sensitive disclosure, data poisoning, excessive agency, and other LLM-specific risks. Use when building chatbots, RAG systems, AI agents, or any application calling LLM APIs to identify relevant vulnerabilities and apply secure patterns.

Generated from the current SKILL.md.

Does this skill cover prompt injection attacks?
Yes. Prompt Injection (LLM01) is a critical category with dedicated rules for preventing both direct and indirect prompt manipulation in chatbots, RAG systems, and tool-using LLMs.
What LLM APIs and libraries does this apply to?
This skill applies to any application calling OpenAI, Anthropic, Gemini, or similar LLM APIs, and to code using LangChain, LlamaIndex, or comparable LLM frameworks.
Should I use this skill only when the user explicitly asks about security?
No. The skill is designed for proactive use: automatically check for relevant security risks whenever building or reviewing LLM applications, chatbots, RAG pipelines, or AI agents — regardless of whether the user mentions security.
Does this cover RAG system security?
Yes. RAG systems have priority rules for Vector/Embedding Weaknesses (LLM08), Prompt Injection (LLM01), Sensitive Disclosure (LLM02), and Misinformation (LLM09).
Are there code examples for each vulnerability?
Yes. Each of the 10 OWASP categories has a dedicated rule file in `rules/` with vulnerable and secure code examples.

Generated from the current SKILL.md. These answers refresh after source changes.