All skills
apple avatar

/model-compression-exploration

@b1cb71b
by appleapple/coreai-models2.2k stars
219

Systematically explore weight compression configurations (quantization and palettization) for a PyTorch model using coreai-opt, presenting a comprehensive overview of accuracy-vs-size tradeoff options. Use this skill whenever the user wants to compress a model, explore quantization or palettization options, understand compression config tradeoffs, reduce model size, or compare different compression techniques. Also trigger when the user mentions coreai-opt compression, weight quantization exploration, or palettization exploration — even if they don't say "explore" explicitly.

Use this Skill: https://skilld.dev/gh/apple/coreai-models/model-compression-exploration

This session only. Nothing lands on disk.

referencesoutput_report.md

≈1.8k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Output Report for Model Compression exploration

The examples below show the table format (column order, Unicode box style, value formatting). The number of rows is governed by SKILL.md's selection rule: exactly 5 representative configs per group spanning the accuracy-vs-size tradeoff, picked from the JSONL after filtering errors and configs below the quality floor. Some examples below show fewer than 5 rows — that's only because the snippet was abbreviated.

Example output report:

  Group 1: Per-Channel Quantization

  ┌────────────────────────────────┬───────────┬──────────────┬───────────────────┐
  │             Config             │ PSNR (dB) │ Avg Bitwidth │ Compression Ratio │
  ├────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ perchannel_int8_symmetric      │ 57.59     │ 8.04         │ 1.99x             │
  ├────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ perchannel_int4_symmetric_skip │ 31.61     │ 4.23         │ 3.78x             │
  ├────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ perchannel_int4_symmetric      │ 30.41     │ 4.04         │ 3.96x             │
  └────────────────────────────────┴───────────┴──────────────┴───────────────────┘


  Group 2: Per-Block Quantization

  ┌───────────────────────────────────┬───────────┬──────────────┬───────────────────┐
  │              Config               │ PSNR (dB) │ Avg Bitwidth │ Compression Ratio │
  ├───────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ perblock_bs32_int8_symmetric      │ 62.71     │ 9.00         │ 1.78x             │
  ├───────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ perblock_bs64_int8_symmetric      │ 62.36     │ 8.50         │ 1.88x             │
  ├───────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ perblock_bs128_int8_symmetric     │ 61.31     │ 8.25         │ 1.94x             │
  ├───────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ perblock_bs64_int4_symmetric_skip │ 36.03     │ 4.68         │ 3.42x             │
  ├───────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ perblock_bs64_int4_symmetric      │ 35.33     │ 4.50         │ 3.56x             │
  └───────────────────────────────────┴───────────┴──────────────┴───────────────────┘

  Group 3: Palettization

  ┌──────────────────────────────────┬───────────┬──────────────┬───────────────────┐
  │              Config              │ PSNR (dB) │ Avg Bitwidth │ Compression Ratio │
  ├──────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ palette_pertensor_8bit           │ 59.39     │ 8.01         │ 2.00x             │
  ├──────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ palette_perchannel_4bit_pcs_skip │ 40.77     │ 4.54         │ 3.53x             │
  ├──────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ palette_perchannel_4bit_skip     │ 40.36     │ 4.52         │ 3.54x             │
  ├──────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ palette_grouped_gs4_4bit         │ 38.95     │ 4.08         │ 3.92x             │
  ├──────────────────────────────────┼───────────┼──────────────┼───────────────────┤
  │ palette_perchannel_4bit_pcs      │ 38.63     │ 4.35         │ 3.67x             │
  └──────────────────────────────────┴───────────┴──────────────┴───────────────────┘

When including configs that have layer skipping, mention the skipped layer type(s) or layer name(s)

Source: SKILL.md on GitHub

No alerts3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill provides a systematic framework for exploring model compression tradeoffs using standard PyTorch workflows. It is authored by a trusted vendor and follows secure patterns for model optimization and data handling.

  • Socket3mo

    No alerts

  • Snyk3mo

    Risk: LOW · No issues

Signed by skilld at b1cb71b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 5 hours ago.

Activeupdated 4 months ago

README badge

README badge for apple/coreai-models/model-compression-exploration