Skip to main content

Common Tasks

3 min read

Control AI Costs

Prevent runaway agents from burning through your API budget.


The Problem

An agent enters a reasoning loop and burns through $50 of tokens in minutes. A batch job scales to 1,000 concurrent agents overnight. Without guardrails on spending, LLM costs spiral out of control.

The Solution

Use maxTokenBudget on the orchestrator for simple token caps, or withBudget middleware for dollar-based limits with time windows:

import {
  createAgentOrchestrator,
  requireModelPricing,
  withBudget,
} from '@directive-run/ai';
import { ANTHROPIC_PRICING } from '@directive-run/ai/anthropic';

// Option 1: Simple token cap on orchestrator
const orchestrator = createAgentOrchestrator({
  runner, // See Running Agents (/ai/running-agents) for setup
  autoApproveToolCalls: true,
  maxTokenBudget: 50000,
  budgetWarningThreshold: 0.8,
  onBudgetWarning: ({ currentTokens, maxBudget, percentage }) => {
    console.warn(`Budget ${Math.round(percentage * 100)}% used`);
  },
});

// Option 2: Dollar-based limits with time windows
// One set of rates, used everywhere. Every window and the top-level `pricing`
// must price a call identically — mismatched rates now throw at construction.
const pricing = requireModelPricing(ANTHROPIC_PRICING, 'claude-sonnet-4-5-20250929');

const budgetedRunner = withBudget(runner, {
  maxCostPerCall: 0.50,
  budgets: [
    { window: 'hour', maxCost: 10.00, pricing },
    { window: 'day', maxCost: 100.00, pricing },
  ],
  pricing,
  onBudgetExceeded: (details) => {
    console.error(
      `Budget exceeded (${details.window}, ${details.phase}): ` +
      `$${details.estimated.toFixed(4)} against $${details.remaining.toFixed(4)} left`
    );
  },
});

How It Works

  • maxTokenBudget is a hard cap on total tokens (input + output) across all agent turns. The agent pauses when exceeded.
  • budgetWarningThreshold fires onBudgetWarning at the specified percentage (0.8 = 80%).
  • withBudget wraps a runner with dollar-based cost tracking. It estimates costs before each call and rejects calls that would exceed the budget.
  • budgets array supports multiple time windows. Each window tracks spending independently.
  • pricing maps token counts to dollar costs, per million tokensinputPerMillion: 3 is $3 per million input tokens. Rather than writing rates by hand, take them from the adapter's table: requireModelPricing(ANTHROPIC_PRICING, 'claude-sonnet-4-5-20250929') returns an entry that works here directly, and throws naming the model if you mistype it. If your provider reports cached tokens, set cacheReadPerMillion and cacheWritePerMillion too — they are billed separately and are not included in the input count.

Full Example

An orchestrator with both token caps and dollar budgets, plus logging:

import {
  createAgentOrchestrator,
  requireModelPricing,
  withBudget,
} from '@directive-run/ai';
import { ANTHROPIC_PRICING } from '@directive-run/ai/anthropic';

const pricing = requireModelPricing(ANTHROPIC_PRICING, 'claude-sonnet-4-5-20250929');

const budgetedRunner = withBudget(runner, { // See Running Agents (/ai/running-agents) for setup
  maxCostPerCall: 1.00,
  budgets: [
    { window: 'hour', maxCost: 25.00, pricing },
    { window: 'day', maxCost: 200.00, pricing },
  ],
  pricing,
  onBudgetExceeded: (details) => {
    // Your alerting function
    alertOps(`AI budget exceeded (${details.window}, ${details.phase}): $${details.estimated.toFixed(4)}`);
  },
});

const orchestrator = createAgentOrchestrator({
  runner: budgetedRunner,
  autoApproveToolCalls: true,
  maxTokenBudget: 100000,
  budgetWarningThreshold: 0.9,
  onBudgetWarning: ({ percentage }) => {
    console.warn(`Token budget at ${Math.round(percentage * 100)}%`);
  },
});

// Check spending at any time
const hourlySpend = budgetedRunner.getSpent('hour');
const dailySpend = budgetedRunner.getSpent('day');
console.log(`Spent: $${hourlySpend.toFixed(2)}/hr, $${dailySpend.toFixed(2)}/day`);
Previous
Handle Agent Errors

Stay in the loop. Sign up for our newsletter.

We care about your data. We'll never share your email.

Powered by Directive. This signup uses a Directive module with facts, derivations, constraints, and resolvers – zero useState, zero useEffect. Read how it works

Directive - Constraint-Driven Runtime for TypeScript | AI Guardrails & State Management