Cost-Aware LLM Pipeline
Original work by ECC. This guide keeps and builds on ECC's core design.
Use these steps to lower LLM API cost while keeping good results:
- Pick a low-cost model for simple work.
- Check the budget before each call.
- Retry only short-term errors.
- Cache long text that does not change.
- Save the real cost after each call.
Do not add tracking, analytics, or calls that are not needed for the task.
Model choice
Keep model names in settings. Do not spread them through the code.
Start with the low-cost model. Use the strong model when the task is large or hard. Text size alone is not always enough. Also check the task type, item count, file count, and need for deep thought.
MODEL_FAST = "your-fast-model"
MODEL_STRONG = "your-strong-model"
STRONG_TEXT_LIMIT = 10_000
STRONG_ITEM_LIMIT = 30
def select_model(
text_length: int,
item_count: int,
needs_deep_work: bool = False,
force_model: str | None = None,
) -> str:
if text_length < 0 or item_count < 0:
raise ValueError("Counts must not be below zero.")
if force_model is not None:
return force_model
if (
needs_deep_work
or text_length >= STRONG_TEXT_LIMIT
or item_count >= STRONG_ITEM_LIMIT
):
return MODEL_STRONG
return MODEL_FASTA forced model must still pass the budget check.
Test the limits with real tasks. Raise them if the fast model works well. Lower them if it fails too often.
Budget checks
Keep cost data unchanged after it is made. Each call returns a new tracker.
Use Decimal for money. A float can cause small math errors.
from dataclasses import dataclass
from decimal import Decimal
@dataclass(frozen=True, slots=True)
class CostRecord:
model: str
input_tokens: int
output_tokens: int
cost_usd: Decimal
@dataclass(frozen=True, slots=True)
class CostTracker:
budget_limit: Decimal
records: tuple[CostRecord, ...] = ()
def add(self, record: CostRecord) -> "CostTracker":
if record.cost_usd < 0:
raise ValueError("Cost must not be below zero.")
return CostTracker(
budget_limit=self.budget_limit,
records=(*self.records, record),
)
@property
def total_cost(self) -> Decimal:
return sum(
(record.cost_usd for record in self.records),
start=Decimal("0"),
)
def can_spend(self, expected_cost: Decimal) -> bool:
if expected_cost < 0:
raise ValueError("Expected cost must not be below zero.")
return self.total_cost + expected_cost <= self.budget_limitCheck the likely cost before the call. A check made only after the call can go over the budget.
Keep a small safety pad for token counts that may be wrong. Stop new work when the next call may pass the limit. Let the caller choose whether to skip the item, wait, or end the batch.
If the API does not return token use, mark the cost as an estimate. Do not record it as exact.
Safe retries
Retry only errors that may go away soon. Examples are lost links, rate limits, and server faults.
Do not retry bad keys, blocked access, bad input, missing models, or content policy errors.
Use delay hints from the API when they exist. Else, wait longer after each failed try. Add a small random wait so many workers do not retry at once.
import random
import time
from collections.abc import Callable
from typing import TypeVar
from anthropic import APIConnectionError, InternalServerError, RateLimitError
T = TypeVar("T")
RETRYABLE_ERRORS = (
APIConnectionError,
RateLimitError,
InternalServerError,
)
def call_with_retry(
func: Callable[[], T],
*,
max_attempts: int = 3,
) -> T:
if max_attempts < 1:
raise ValueError("max_attempts must be at least 1.")
for attempt in range(max_attempts):
try:
return func()
except RETRYABLE_ERRORS:
if attempt + 1 == max_attempts:
raise
delay = (2 ** attempt) + random.uniform(0, 0.5)
time.sleep(delay)
raise RuntimeError("Retry loop ended without a result.")Set request time limits. A stuck call should not wait forever.
A retry can create a second result if the first call worked but its reply was lost. Use a request key when the API supports one. Make writes safe to run twice.
Prompt caching
Cache long text that stays the same across calls. Keep user text and other changing text outside the cached block.
def build_messages(rules: str, user_text: str) -> list[dict]:
return [
{
"role": "user",
"content": [
{
"type": "text",
"text": rules,
"cache_control": {"type": "ephemeral"},
},
{
"type": "text",
"text": user_text,
},
],
}
]Follow the cache form for the API and model you use. Cache rules differ by vendor. Some models do not support caching.
Do not place private user data in a shared cache. Do not assume a cache hit. The request must still work after a cache miss or cache end.
Full pipeline
Check the likely cost first. Call the model. Then save the real cost from the reply.
from dataclasses import dataclass
from decimal import Decimal
class BudgetExceededError(RuntimeError):
pass
@dataclass(frozen=True, slots=True)
class Config:
force_model: str | None = None
max_attempts: int = 3
def process(
text: str,
item_count: int,
rules: str,
config: Config,
tracker: CostTracker,
):
model = select_model(
text_length=len(text),
item_count=item_count,
force_model=config.force_model,
)
expected_cost = estimate_cost(
model=model,
text=text,
rules=rules,
)
if not tracker.can_spend(expected_cost):
raise BudgetExceededError(
f"Next call may pass the ${tracker.budget_limit} budget."
)
response = call_with_retry(
lambda: client.messages.create(
model=model,
max_tokens=800,
messages=build_messages(rules, text),
),
max_attempts=config.max_attempts,
)
usage = response.usage
actual_cost = price_usage(
model=model,
input_tokens=usage.input_tokens,
output_tokens=usage.output_tokens,
)
tracker = tracker.add(
CostRecord(
model=model,
input_tokens=usage.input_tokens,
output_tokens=usage.output_tokens,
cost_usd=actual_cost,
)
)
return parse_result(response), trackerestimate_cost, price_usage, client, and parse_result are app parts. Keep price data in one settings file. Use the model ID returned by the API when possible.
Example
This job has a $1 budget. It stops before the next call may pass that limit.
from decimal import Decimal
tracker = CostTracker(budget_limit=Decimal("1.00"))
config = Config(max_attempts=3)
items = [
"Give this note a short title.",
"Compare these 40 reports and list the main risks.",
]
for text in items:
try:
result, tracker = process(
text=text,
item_count=1,
rules="Use plain words. Return valid JSON.",
config=config,
tracker=tracker,
)
print(result)
except BudgetExceededError:
print("Budget reached. The rest of the batch was skipped.")
break
print(f"Spent: ${tracker.total_cost}")For a real batch, pass the true item count. Set needs_deep_work=True for work that is short but hard.
Price data
Do not copy old prices into code. Model names and prices change.
Keep one checked price table in settings. Record:
- The full model ID
- Input price
- Output price
- Cache write price, if used
- Cache read price, if used
- The date the price was checked
Round only when you show a total. Keep full value while you add costs.
Checks before release
- Test a simple task and a hard task.
- Test a task on each model limit.
- Test a budget equal to the next call cost.
- Test a budget just below the next call cost.
- Test one short-term error, then success.
- Test all retry tries failing.
- Test a bad key with no retry.
- Test a cache miss and a cache hit.
- Test a reply with missing token data.
- Test more than one worker using the same budget.
If many workers share one budget, store the total in one safe place. Lock the check and update as one step. A local tracker alone cannot stop two workers from spending the same money.
Rules
- Start with the lowest-cost model that can do the work.
- Let the user force a model, but never skip the budget check.
- Set a firm output token limit.
- Check likely cost before each call.
- Save real token use and cost after each call.
- Retry only short-term errors.
- Cap retry tries and wait time.
- Keep price data and model names in one place.
- Do not send logs, cost data, or task data to any extra service.