All skills

Build LLM pipelines that control cost with model choice, budget checks, safe retries, and prompt caching. Use for batch jobs, mixed task sizes, or any app with a firm API budget.

  • 1 file
  • 9.2 KB
  • Updated 4 weeks ago
  • GitHub

Use this Skill: https://skilld.dev/gh/agenticluke/llm-cost-guard-plus/skill

This session only. Nothing lands on disk.

SKILL.md

≈46 tokens always: the name and description. ≈2.3k when used: this file.

Cost-Aware LLM Pipeline

Original work by ECC. This guide keeps and builds on ECC's core design.

Use these steps to lower LLM API cost while keeping good results:

  1. Pick a low-cost model for simple work.
  2. Check the budget before each call.
  3. Retry only short-term errors.
  4. Cache long text that does not change.
  5. Save the real cost after each call.

Do not add tracking, analytics, or calls that are not needed for the task.

Model choice

Keep model names in settings. Do not spread them through the code.

Start with the low-cost model. Use the strong model when the task is large or hard. Text size alone is not always enough. Also check the task type, item count, file count, and need for deep thought.

MODEL_FAST = "your-fast-model"
MODEL_STRONG = "your-strong-model"

STRONG_TEXT_LIMIT = 10_000
STRONG_ITEM_LIMIT = 30

def select_model(
    text_length: int,
    item_count: int,
    needs_deep_work: bool = False,
    force_model: str | None = None,
) -> str:
    if text_length < 0 or item_count < 0:
        raise ValueError("Counts must not be below zero.")

    if force_model is not None:
        return force_model

    if (
        needs_deep_work
        or text_length >= STRONG_TEXT_LIMIT
        or item_count >= STRONG_ITEM_LIMIT
    ):
        return MODEL_STRONG

    return MODEL_FAST

A forced model must still pass the budget check.

Test the limits with real tasks. Raise them if the fast model works well. Lower them if it fails too often.

Budget checks

Keep cost data unchanged after it is made. Each call returns a new tracker.

Use Decimal for money. A float can cause small math errors.

from dataclasses import dataclass
from decimal import Decimal

@dataclass(frozen=True, slots=True)
class CostRecord:
    model: str
    input_tokens: int
    output_tokens: int
    cost_usd: Decimal

@dataclass(frozen=True, slots=True)
class CostTracker:
    budget_limit: Decimal
    records: tuple[CostRecord, ...] = ()

    def add(self, record: CostRecord) -> "CostTracker":
        if record.cost_usd < 0:
            raise ValueError("Cost must not be below zero.")

        return CostTracker(
            budget_limit=self.budget_limit,
            records=(*self.records, record),
        )

    @property
    def total_cost(self) -> Decimal:
        return sum(
            (record.cost_usd for record in self.records),
            start=Decimal("0"),
        )

    def can_spend(self, expected_cost: Decimal) -> bool:
        if expected_cost < 0:
            raise ValueError("Expected cost must not be below zero.")

        return self.total_cost + expected_cost <= self.budget_limit

Check the likely cost before the call. A check made only after the call can go over the budget.

Keep a small safety pad for token counts that may be wrong. Stop new work when the next call may pass the limit. Let the caller choose whether to skip the item, wait, or end the batch.

If the API does not return token use, mark the cost as an estimate. Do not record it as exact.

Safe retries

Retry only errors that may go away soon. Examples are lost links, rate limits, and server faults.

Do not retry bad keys, blocked access, bad input, missing models, or content policy errors.

Use delay hints from the API when they exist. Else, wait longer after each failed try. Add a small random wait so many workers do not retry at once.

import random
import time
from collections.abc import Callable
from typing import TypeVar

from anthropic import APIConnectionError, InternalServerError, RateLimitError

T = TypeVar("T")

RETRYABLE_ERRORS = (
    APIConnectionError,
    RateLimitError,
    InternalServerError,
)

def call_with_retry(
    func: Callable[[], T],
    *,
    max_attempts: int = 3,
) -> T:
    if max_attempts < 1:
        raise ValueError("max_attempts must be at least 1.")

    for attempt in range(max_attempts):
        try:
            return func()
        except RETRYABLE_ERRORS:
            if attempt + 1 == max_attempts:
                raise

            delay = (2 ** attempt) + random.uniform(0, 0.5)
            time.sleep(delay)

    raise RuntimeError("Retry loop ended without a result.")

Set request time limits. A stuck call should not wait forever.

A retry can create a second result if the first call worked but its reply was lost. Use a request key when the API supports one. Make writes safe to run twice.

Prompt caching

Cache long text that stays the same across calls. Keep user text and other changing text outside the cached block.

def build_messages(rules: str, user_text: str) -> list[dict]:
    return [
        {
            "role": "user",
            "content": [
                {
                    "type": "text",
                    "text": rules,
                    "cache_control": {"type": "ephemeral"},
                },
                {
                    "type": "text",
                    "text": user_text,
                },
            ],
        }
    ]

Follow the cache form for the API and model you use. Cache rules differ by vendor. Some models do not support caching.

Do not place private user data in a shared cache. Do not assume a cache hit. The request must still work after a cache miss or cache end.

Full pipeline

Check the likely cost first. Call the model. Then save the real cost from the reply.

from dataclasses import dataclass
from decimal import Decimal

class BudgetExceededError(RuntimeError):
    pass

@dataclass(frozen=True, slots=True)
class Config:
    force_model: str | None = None
    max_attempts: int = 3

def process(
    text: str,
    item_count: int,
    rules: str,
    config: Config,
    tracker: CostTracker,
):
    model = select_model(
        text_length=len(text),
        item_count=item_count,
        force_model=config.force_model,
    )

    expected_cost = estimate_cost(
        model=model,
        text=text,
        rules=rules,
    )

    if not tracker.can_spend(expected_cost):
        raise BudgetExceededError(
            f"Next call may pass the ${tracker.budget_limit} budget."
        )

    response = call_with_retry(
        lambda: client.messages.create(
            model=model,
            max_tokens=800,
            messages=build_messages(rules, text),
        ),
        max_attempts=config.max_attempts,
    )

    usage = response.usage
    actual_cost = price_usage(
        model=model,
        input_tokens=usage.input_tokens,
        output_tokens=usage.output_tokens,
    )

    tracker = tracker.add(
        CostRecord(
            model=model,
            input_tokens=usage.input_tokens,
            output_tokens=usage.output_tokens,
            cost_usd=actual_cost,
        )
    )

    return parse_result(response), tracker

estimate_cost, price_usage, client, and parse_result are app parts. Keep price data in one settings file. Use the model ID returned by the API when possible.

Example

This job has a $1 budget. It stops before the next call may pass that limit.

from decimal import Decimal

tracker = CostTracker(budget_limit=Decimal("1.00"))
config = Config(max_attempts=3)

items = [
    "Give this note a short title.",
    "Compare these 40 reports and list the main risks.",
]

for text in items:
    try:
        result, tracker = process(
            text=text,
            item_count=1,
            rules="Use plain words. Return valid JSON.",
            config=config,
            tracker=tracker,
        )
        print(result)
    except BudgetExceededError:
        print("Budget reached. The rest of the batch was skipped.")
        break

print(f"Spent: ${tracker.total_cost}")

For a real batch, pass the true item count. Set needs_deep_work=True for work that is short but hard.

Price data

Do not copy old prices into code. Model names and prices change.

Keep one checked price table in settings. Record:

  • The full model ID
  • Input price
  • Output price
  • Cache write price, if used
  • Cache read price, if used
  • The date the price was checked

Round only when you show a total. Keep full value while you add costs.

Checks before release

  • Test a simple task and a hard task.
  • Test a task on each model limit.
  • Test a budget equal to the next call cost.
  • Test a budget just below the next call cost.
  • Test one short-term error, then success.
  • Test all retry tries failing.
  • Test a bad key with no retry.
  • Test a cache miss and a cache hit.
  • Test a reply with missing token data.
  • Test more than one worker using the same budget.

If many workers share one budget, store the total in one safe place. Lock the check and update as one step. A local tracker alone cannot stop two workers from spending the same money.

Rules

  • Start with the lowest-cost model that can do the work.
  • Let the user force a model, but never skip the budget check.
  • Set a firm output token limit.
  • Check likely cost before each call.
  • Save real token use and cost after each call.
  • Retry only short-term errors.
  • Cap retry tries and wait time.
  • Keep price data and model names in one place.
  • Do not send logs, cost data, or task data to any extra service.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at f08bb45. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 4 weeks ago.

Activeupdated 4 weeks ago
origin
ECC

README badge

README badge for agenticluke/llm-cost-guard-plus