All skills
nimrodfisher avatar

/programmatic-eda

@2e18ac4

Systematic exploratory data analysis. Activate when a dataset needs profiling — structure check, nulls, outliers, distributions, correlations — before deeper analysis begins.

Use this Skill: https://skilld.dev/gh/nimrodfisher/data-analytics-skills/programmatic-eda

This session only. Nothing lands on disk.

SKILL.md

≈49 tokens always: the name and description. ≈532 when used: this file. ≈3.3k more on demand in 5 files.

When to use

  • You receive a new dataset and need to understand its shape and quality before analysis
  • An analysis produces surprising numbers and you want to verify the underlying data first
  • A stakeholder asks "is this data reliable?" or "what's in this table?"
  • You're about to run a model or statistical test and need data-quality assurance

Process

  1. Load and overview — run scripts/data_overview.py to get row count, dtypes, memory usage, and a sample. Confirm grain (what one row represents).
  2. Null profile — run scripts/null_profiler.py; compare output against thresholds in references/quality_thresholds.md and flag columns above limits.
  3. Outlier detection — run scripts/outlier_detector.py (IQR + z-score) on numeric columns; document flagged values and decide: real signal or data error?
  4. Distribution summary — run scripts/distribution_summary.py for descriptive stats and univariate histograms on each numeric column.
  5. Correlation exploration — run scripts/correlation_explorer.py; flag pairs with |r| > 0.8 as potential multicollinearity or redundancy.
  6. EDA checklist sign-off — work through references/eda_checklist.md and confirm each item before declaring the dataset profiled.
  7. Write findings — fill assets/eda_report_template.md with full profiling output; distil top issues into assets/findings_summary.md.

For pattern recipes (e.g. polars vs pandas equivalents, chunked reads for large files), see references/pandas_polars_recipes.md.

Inputs the skill needs

  • Required: dataset path (CSV / Parquet / Excel) or a DataFrame already in scope
  • Required: business context — what does one row represent?
  • Optional: quality threshold overrides (defaults in references/quality_thresholds.md)
  • Optional: columns to skip (PII, binary blobs, high-cardinality IDs)

Output

  • assets/eda_report_template.md (filled) — full profiling report with per-column stats
  • assets/findings_summary.md (filled) — top 3–5 quality issues and recommended next steps
  • Console output / plots from scripts for interactive inspection

Source: SKILL.md on GitHub

No alerts16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The skill provides a set of Python scripts and markdown templates for performing systematic exploratory data analysis (EDA) on local datasets. It performs data profiling, statistical analysis, and report generation locally without any detected malicious patterns or dangerous network operations.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 2e18ac4. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 5 days ago.

Activeupdated 5 months ago

README badge

README badge for nimrodfisher/data-analytics-skills/programmatic-eda