All skills
hardw00t avatar

/sast-orchestration

@f9bb3b2

Static Application Security Testing orchestration — run and compose Semgrep, CodeQL, Bandit, gosec, Brakeman, SpotBugs, ESLint; author custom rules; ingest SARIF; triage and rank findings by exploitability. Use this skill when asked to scan code for vulnerabilities, write Semgrep/CodeQL rules, triage SAST output, reduce false positives, or integrate SAST into CI/CD. Triggers on phrases like 'scan this code', 'write a Semgrep rule', 'triage these findings', 'SARIF', 'SAST in CI', or when a repo is handed over for a security review.

Use this Skill: https://skilld.dev/gh/hardw00t/ai-security-arsenal/sast-orchestration

This session only. Nothing lands on disk.

workflowscustom_codeql_from_sink.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Workflow: CodeQL Taint-Tracking Query from a Sink

Given a new dangerous API (sink) — e.g., myorm.raw_query(sql), internal.exec(cmd) — write a taint-tracking query that flags user input reaching it.

Reasoning budget: HIGH

CodeQL query design requires:

  • Knowing the right standard-library module for sources (flask, express, servlet).
  • Modeling the sink shape in QL (method? constructor? argument index?).
  • Choosing sanitizers to silence known-safe paths.
  • Deciding between problem and path-problem (path-problem is almost always right for taint).

Budget: expect multiple iterations. Use the CodeQL extension in VS Code to iterate.

Inputs

  • Sink API: fully-qualified name, argument that takes the dangerous value, sanitizer names (if any).
  • Language.
  • Existing CodeQL database (or create one — see references/codeql.md).

Template (Python)

Use examples/codeql_queries/taint_template.ql as the starting point. Adapt the three predicates: isSource, isSink, isBarrier.

/**
 * @name Custom taint to myorm.raw_query
 * @description Untrusted input reaches myorm.raw_query
 * @kind path-problem
 * @problem.severity error
 * @security-severity 8.8
 * @precision medium
 * @id py/custom-taint-myorm-raw
 * @tags security external/cwe/cwe-089
 */

import python
import semmle.python.dataflow.new.TaintTracking
import semmle.python.dataflow.new.DataFlow
import semmle.python.ApiGraphs

module CustomConfig implements DataFlow::ConfigSig {
  predicate isSource(DataFlow::Node src) {
    // HTTP inputs — Flask, Django, FastAPI shapes
    src = API::moduleImport("flask").getMember("request").getMember(_).getACall()
    or
    src = API::moduleImport("django").getMember("http").getMember("HttpRequest").getInstance().getMember(_).getACall()
    or
    exists(API::Node fastapi |
      fastapi = API::moduleImport("fastapi") and
      src = fastapi.getMember("Request").getInstance().getMember(_).getACall()
    )
  }

  predicate isSink(DataFlow::Node sink) {
    exists(DataFlow::CallCfgNode c |
      c = API::moduleImport("myorm").getMember("raw_query").getACall() and
      sink = c.getArg(0)
    )
  }

  predicate isBarrier(DataFlow::Node n) {
    // Known safe: parameterized wrapper
    n = API::moduleImport("myorm").getMember("quote_literal").getACall()
  }
}

module CustomFlow = TaintTracking::Global<CustomConfig>;
import CustomFlow::PathGraph

from CustomFlow::PathNode source, CustomFlow::PathNode sink
where CustomFlow::flowPath(source, sink)
select sink.getNode(), source, sink,
       "User input from $@ flows to myorm.raw_query.",
       source.getNode(), "HTTP request"

Language variants

JavaScript / TypeScript

import javascript
import semmle.javascript.security.dataflow.TaintTracking

module CustomConfig implements DataFlow::ConfigSig {
  predicate isSource(DataFlow::Node src) {
    src instanceof RemoteFlowSource  // built-in: req params, body, headers
  }
  predicate isSink(DataFlow::Node sink) {
    exists(DataFlow::CallNode c |
      c = API::moduleImport("myorm").getMember("rawQuery").getACall() and
      sink = c.getArgument(0)
    )
  }
}
module CustomFlow = TaintTracking::Global<CustomConfig>;

Java

import java
import semmle.code.java.dataflow.FlowSources
import semmle.code.java.dataflow.TaintTracking

module CustomConfig implements DataFlow::ConfigSig {
  predicate isSource(DataFlow::Node src) {
    src instanceof RemoteFlowSource
  }
  predicate isSink(DataFlow::Node sink) {
    exists(MethodCall mc |
      mc.getMethod().hasQualifiedName("com.acme.myorm", "MyOrm", "rawQuery") and
      sink.asExpr() = mc.getArgument(0)
    )
  }
}
module CustomFlow = TaintTracking::Global<CustomConfig>;

Go

import go
import semmle.go.dataflow.TaintTracking

module CustomConfig implements DataFlow::ConfigSig {
  predicate isSource(DataFlow::Node src) { src instanceof UntrustedFlowSource::Range }
  predicate isSink(DataFlow::Node sink) {
    exists(DataFlow::CallNode c |
      c.getTarget().hasQualifiedName("github.com/acme/myorm", "RawQuery") and
      sink = c.getArgument(0)
    )
  }
}
module CustomFlow = TaintTracking::Global<CustomConfig>;

Sink-shape reference

Sink shape QL snippet
Function top-level API::moduleImport("pkg").getMember("fn").getACall()
Method on class instance API::moduleImport("pkg").getMember("Cls").getInstance().getMember("m").getACall()
Constructor argument API::moduleImport("pkg").getMember("Cls").getACall() (the call itself is the instance)
Nth positional arg .getArg(n) (Py) / .getArgument(n) (JS/Java)
Keyword arg .getArgByName("key")

Sources cheat sheet

Language Built-in "untrusted" source
Python Manual: flask.request.*, django.http.HttpRequest.*, fastapi.Request.*
JavaScript RemoteFlowSource
Java RemoteFlowSource
Go UntrustedFlowSource::Range
Ruby Http::ActiveRecordSqlExecutionRange et al — check codeql/ruby-queries
C/C++ FlowSource (manual modeling required)

Iteration loop

  1. Write minimal isSource + isSink with an empty isBarrier.
  2. Run against a test DB seeded with a known TP and a known TN.
  3. If FN on TP: check sink shape (usually argument index wrong).
  4. If FP on TN: add isBarrier for the sanitizer used there.
  5. Raise @precision once FP rate stabilizes <10%.

Test harness

Place a .ql file next to a .expected file with expected results. Use codeql test run <path>.

When to use CodeQL vs Semgrep for custom rules

  • If inter-procedural flow across >2 files is required → CodeQL.
  • If the sink is one well-defined API and codebase already has a CodeQL DB → CodeQL.
  • If rule must run in every PR on every language → Semgrep (faster iteration).
  • If you need high precision for a single vulnerability family → CodeQL.

Source: SKILL.md on GitHub

1 warning16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill is a security orchestration suite designed to run and aggregate results from various Static Application Security Testing (SAST) tools. It includes workflows for scanning codebases, triaging findings, and authoring custom detection rules. No malicious patterns or security risks were identified; the skill correctly manages its capabilities to provide a comprehensive security analysis environment.

  • Socket16d

    1 alert: gptSecurity

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer7mo

    1 file scanned · No issues

Signed by skilld at f9bb3b2. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Steadyupdated 6 months ago

README badge

README badge for hardw00t/ai-security-arsenal/sast-orchestration