All skills
huggingface avatar

/transformers-js

@4db6015 official
by Hugging Facehuggingface/skills11k stars
753

Use Transformers.js to run state-of-the-art machine learning models directly in JavaScript/TypeScript. Supports NLP (text classification, translation, summarization), computer vision (image classification, object detection), audio (speech recognition, audio classification), and multimodal tasks. Works in browsers and server-side runtimes (Node.js, Bun, Deno) with WebGPU/WASM using pre-trained models from Hugging Face Hub.

Use this Skill: https://skilld.dev/gh/huggingface/skills/transformers-js

This session only. Nothing lands on disk.

referencesMODEL_REGISTRY.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

ModelRegistry Reference

In Transformers.js v4, ModelRegistry provides a preflight API for model assets. You can inspect required files, estimate total download size, check cache state, and clear cached artifacts before calling pipeline().

This is useful for production UX where you want to:

  • show accurate download estimates before loading,
  • support offline-first flows,
  • avoid surprise bandwidth usage,
  • and keep cache management explicit.

Table of Contents

  1. Overview
  2. Core APIs
  3. Recommended Workflow
  4. Examples
  5. Best Practices

Overview

import { ModelRegistry } from '@huggingface/transformers';

ModelRegistry works with the same task/model/options you pass to pipeline().

Typical tuple:

const task = 'feature-extraction';
const modelId = 'onnx-community/all-MiniLM-L6-v2-ONNX';
const modelOptions = { dtype: 'fp32' };

Core APIs

get_pipeline_files(task, modelId, modelOptions)

Returns all files needed to initialize that pipeline configuration.

const files = await ModelRegistry.get_pipeline_files(task, modelId, modelOptions);
// Example: ['config.json', 'onnx/model.onnx', 'tokenizer.json', ...]

Use this to build preflight checks and download manifests.

get_file_metadata(modelId, file)

Returns metadata for a single file (including size when available).

const metadata = await ModelRegistry.get_file_metadata(modelId, 'onnx/model.onnx');
console.log(metadata);

Use this to compute total transfer size and identify large artifacts.

is_pipeline_cached(task, modelId, modelOptions)

Checks whether required files are already available in cache.

const cached = await ModelRegistry.is_pipeline_cached(task, modelId, modelOptions);
console.log(cached ? 'Ready offline' : 'Needs download');

Use this to gate offline mode and skip unnecessary preload steps.

clear_pipeline_cache(task, modelId, modelOptions)

Clears cached assets for a specific pipeline tuple.

await ModelRegistry.clear_pipeline_cache(task, modelId, modelOptions);

Use this for cache invalidation, testing, or space reclamation.

get_available_dtypes(modelId)

Returns precision/quantization formats available for the model.

const dtypes = await ModelRegistry.get_available_dtypes(modelId);
// Example: ['fp32', 'fp16', 'q4', 'q4f16']

Use this to choose the best runtime profile (quality vs. speed vs. memory).

Recommended Workflow

For robust loading UX:

  1. Resolve the exact task/model/options tuple.
  2. Call get_pipeline_files(...).
  3. Fetch metadata per file and compute total size.
  4. Call is_pipeline_cached(...).
  5. If not cached, show user-facing size/progress expectations.
  6. Load via pipeline(...) and use progress_total in progress_callback.
import { ModelRegistry, pipeline } from '@huggingface/transformers';

const task = 'feature-extraction';
const modelId = 'onnx-community/all-MiniLM-L6-v2-ONNX';
const modelOptions = { dtype: 'q8' };

const files = await ModelRegistry.get_pipeline_files(task, modelId, modelOptions);

const metadata = await Promise.all(
  files.map((file) => ModelRegistry.get_file_metadata(modelId, file))
);

const totalBytes = metadata.reduce((sum, item) => sum + (item?.size ?? 0), 0);
const totalMB = (totalBytes / 1024 / 1024).toFixed(2);

const cached = await ModelRegistry.is_pipeline_cached(task, modelId, modelOptions);
console.log({ fileCount: files.length, totalMB, cached });

const pipe = await pipeline(task, modelId, {
  ...modelOptions,
  progress_callback: (info) => {
    if (info.status === 'progress_total') {
      console.log(`Loading: ${info.progress.toFixed(1)}%`);
    }
  },
});

await pipe.dispose();

Examples

Example 1: Offer dtype choice dynamically

import { ModelRegistry, pipeline } from '@huggingface/transformers';

const task = 'text-generation';
const modelId = 'onnx-community/Qwen2.5-0.5B-Instruct';

const dtypes = await ModelRegistry.get_available_dtypes(modelId);
const preferred = dtypes.includes('q4') ? 'q4' : dtypes[0] ?? 'fp32';

const generator = await pipeline(task, modelId, { dtype: preferred });
// ... inference
await generator.dispose();

Example 2: Only clear one pipeline cache entry

import { ModelRegistry } from '@huggingface/transformers';

await ModelRegistry.clear_pipeline_cache(
  'feature-extraction',
  'onnx-community/all-MiniLM-L6-v2-ONNX',
  { dtype: 'fp32' }
);

This avoids wiping unrelated model caches.

Example 3: Offline gate

import { ModelRegistry, env, pipeline } from '@huggingface/transformers';

const task = 'feature-extraction';
const modelId = 'onnx-community/all-MiniLM-L6-v2-ONNX';
const modelOptions = { dtype: 'q8' };

const cached = await ModelRegistry.is_pipeline_cached(task, modelId, modelOptions);

if (!cached) {
  throw new Error('Model not cached yet. Connect once to download assets.');
}

env.allowRemoteModels = false;
const pipe = await pipeline(task, modelId, { ...modelOptions, local_files_only: true });

Best Practices

  1. Use ModelRegistry before pipeline() when you need predictable download UX.
  2. Cache decisions should be per task/model/options tuple (dtype and revision matter).
  3. Use progress_total for user-facing progress bars; keep per-file progress optional.
  4. Prefer selective invalidation with clear_pipeline_cache(...) over broad cache deletion.
  5. In offline mode, combine is_pipeline_cached(...) with local_files_only: true and env.allowRemoteModels = false.

Related Documentation

Source: SKILL.md on GitHub

No alerts16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides comprehensive documentation and examples for using Transformers.js to run machine learning models. All external dependencies and model resources originate from well-known registries or trusted vendor domains, following industry-standard practices for client-side machine learning.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    5/7 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 4db6015. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 6 months ago
Other metadata
metadata
{
  "author": "huggingface",
  "version": "4.x",
  "category": "machine-learning",
  "repository": "https://github.com/huggingface/transformers.js"
}
compatibility
Requires Node.js 18+ (or compatible Bun/Deno runtime) or modern browser with ES modules support. WebGPU requires runtime and hardware support; WASM is the broad fallback. Internet access is needed for downloading models from Hugging Face Hub (optional if using local models).
  • TypeScript
  • transformers.js
  • machine-learning
  • nlp
  • computer-vision
  • javascript
  • huggingface
  • text-classification
  • object-detection
  • speech-recognition
  • image-classification

README badge

README badge for huggingface/skills/transformers-js

Runs state-of-the-art ML models directly in JavaScript and TypeScript across browsers and Node.js/Bun/Deno using Transformers.js. Supports NLP tasks (text classification, translation, summarization), computer vision (image classification, object detection), audio (speech recognition), and multimodal workflows with models from Hugging Face Hub.

Generated from the current SKILL.md.

Does this work in the browser?
Yes. Transformers.js runs in modern browsers with ES modules support and can be loaded via CDN. Models run client-side using WebGPU (if available) or WASM as a fallback, with no backend server required.
What runtimes does this support?
Node.js 18+, Bun, Deno, and modern browsers. WebGPU requires runtime and hardware support; WASM is the broad fallback for CPU inference.
Do I need to download models ahead of time?
No. Models are downloaded automatically from Hugging Face Hub on first use. Internet access is required unless you configure local models.
What ML tasks does this support?
NLP (text classification, translation, summarization, question-answering, token classification), computer vision (image classification, object detection, segmentation, depth estimation), audio (speech recognition, audio classification, text-to-speech), and multimodal tasks (image captioning, document QA).
Is memory management required?
Yes. All pipelines must be disposed with `pipe.dispose()` when finished to prevent memory leaks.

Generated from the current SKILL.md. These answers refresh after source changes.