All skills
github avatar

/scaling-data-volume

@9637e1a official
by githubgithub/awesome-copilot40k stars
5,040

Guides Qdrant data volume scaling decisions. Use when someone asks 'data doesn't fit on one node', 'too much data', 'need more storage', 'vertical or horizontal scaling', 'tenant scaling', 'time window rotation', or 'data growth exceeds capacity'.

  • 5 files
  • 16.1 KB
  • Updated 6 months ago
  • GitHub

Use this Skill: https://skilld.dev/gh/github/awesome-copilot/scaling-data-volume

This session only. Nothing lands on disk.

SKILL.md

≈67 tokens always: the name and description. ≈399 when used: this file. ≈3.6k more on demand in 4 files.

Scaling Data Volume

This document covers data volume scaling scenarios, where the total size of the dataset exceeds the capacity of a single node.

Tenant Scaling

If the use case is multi-tenant, meaning that each user only has access to a subset of the data, and we never need to query across all the data, then we can use multi-tenancy patterns to scale.

The recommended way is to use multi-tenant workloads with payload partitioning, per-tenant indexes, and tiered multitenancy.

Learn more Tenant Scaling

Sliding Time Window

Some use-cases are based on a sliding time window, where only the most recent data is relevant. For example an index for social media posts, where only the last 6 months of data require fast search.

Learn more Sliding Time Window

Global Search

Most general use-cases require global search across all data. In these situations, we might need to fall back to vertical scaling, and then horizontal scaling when we reach the limits of vertical scaling.

Vertical Scaling

When data doesn't fit in a single node, the first approach is to scale the node itself — more RAM, better disk, quantization, mmap. Exhaust vertical options before going horizontal, as horizontal scaling adds permanent operational complexity.

Learn more Vertical Scaling

Horizontal Scaling

When a single node can't hold the data even with quantization and mmap, distribute data across multiple nodes via sharding.

Learn more Horizontal Scaling

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 9637e1a. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 6 months ago
What it can do
Reads files
All 3 allowed tools
ReadGrepGlob
  • qdrant
  • scaling
  • vector-database
  • sharding
  • multi-tenancy
  • data-volume
  • vertical-scaling
  • horizontal-scaling

README badge

README badge for github/awesome-copilot/scaling-data-volume

Guides decisions for scaling Qdrant vector database when data exceeds single-node capacity. Covers tenant partitioning, time-window rotation, vertical scaling (RAM, quantization, mmap), and horizontal scaling via sharding.

Generated from the current SKILL.md.

When should I use tenant scaling versus horizontal scaling?
Use tenant scaling if each user only accesses a subset of data and you never query across all tenants. Use horizontal scaling for general use-cases that require global search across all data.
Should I scale vertically or horizontally first?
Exhaust vertical scaling options (more RAM, better disk, quantization, mmap) before going horizontal, since horizontal scaling adds permanent operational complexity.
What approach works for time-series or sliding window data?
If only recent data needs fast search (e.g. social media posts from the last 6 months), use sliding time window rotation to manage data growth without scaling the full dataset.

Generated from the current SKILL.md. These answers refresh after source changes.