All skills
nvidia avatar

/data-designer

@e695a83
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use when the user wants to create a dataset, generate synthetic data, or build a data generation pipeline.

Use this Skill: https://skilld.dev/gh/nvidia/skills/data-designer

This session only. Nothing lands on disk.

referencesperson-sampling.md

≈589 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Person Sampling Reference

Sampler types

Prefer "person" when the locale is downloaded — it provides census-grounded demographics and optional personality traits. Fall back to "person_from_faker" when the locale isn't available.

sampler_type Params class When to use
"person" PersonSamplerParams Preferred. Locale downloaded to ~/.data-designer/managed-assets/datasets/ by default.
"person_from_faker" PersonFromFakerSamplerParams Fallback when locale not downloaded. Basic names/addresses via Faker, not demographically accurate.

Usage

The sampled person column is a nested dict. You can keep it as-is in the final dataset, or set drop=True to remove it and extract only the fields you need via ExpressionColumnConfig:

# Keep the full person dict in the output
config_builder.add_column(dd.SamplerColumnConfig(
    name="person", sampler_type="person",
    params=dd.PersonSamplerParams(locale="en_US"),
))

# Or drop it and extract specific fields
config_builder.add_column(dd.SamplerColumnConfig(
    name="person", sampler_type="person",
    params=dd.PersonSamplerParams(locale="en_US"), drop=True,
))
config_builder.add_column(dd.ExpressionColumnConfig(
    name="full_name",
    expr="{{ person.first_name }} {{ person.last_name }}", dtype="str",
))

Set with_synthetic_personas=True when the dataset benefits from personality traits, interests, cultural background, or detailed persona descriptions (e.g., for realistic user simulation or persona-driven prompting). This option is only available with "person" — "person_from_faker" does not support it.

Person Object Schema

Fields vary by locale. Always run the following script to get the exact schema for the locale you are using (script path is relative to this skill's directory):

python scripts/get_person_object_schema.py <locale>

This prints the PII fields (always included) and synthetic persona fields (only included when with_synthetic_personas=True) available for that locale.

Source: SKILL.md on GitHub

No alerts3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    The skill provides a structured framework for generating synthetic datasets using the NVIDIA Data Designer library. It includes workflows for autonomous and interactive design, utilities for inspecting persona schemas, and guidelines for using seed datasets. No security risks were identified beyond standard operational behaviors.

  • Socket3mo

    No alerts

  • Snyk3mo

    Risk: LOW · No issues

Signed by skilld at e695a83. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago
argument-hint
[
  "describe the dataset you want to generate"
]
metadata
{
  "owner": "DataDesigner"
}

README badge

README badge for nvidia/skills/data-designer