Skip to main content
SEP 2026 Latest guide updates. Latest: AI FinOps Token-Saving Tools LLM Market Snapshot Changelog →
3 cost regimes 10 levers Prices verified 2026-09-30

AI FinOps for coding agents

In September 2026, OpenAI announced it would halve a subscription allowance while its API price list offered a model at one fifth of the Astra rate. Which cost regime carries your work decides what changed for you.

September 2026 in three signals

By September 30, 2026, OpenAI's API price list showed a model at one fifth of the Astra rate, and the Pro $200 allowance was set to halve. Anthropic's separate credit for programmatic usage, announced in May, is still paused. None of the three raises a per-token list price.

$2 / $10
GPT-6.1 Sol per 1M tokens, one fifth of GPT-6 Astra
OpenAI model page β†—
20x β†’ 10x
Plus allowance included in ChatGPT Pro $200 from Oct 30, same price
The Next Web, Sep 29, 2026 β†—
Paused
Anthropic's separate credit for Agent SDK and claude -p usage
Claude Help Center β†—

Three cost regimes

Each one has its own unit and moves on its own. Do not convert a quota into tokens: vendors weight allowances by model, effort, and context size.

Subscription

Subscription quota

Unit
A share of an allowance not published in tokens
Examples
  • Claude Pro, Max, Team
  • ChatGPT Plus, Pro
  • Cursor
  • Copilot
What moves it
  • Plan multipliers
  • 5-hour and weekly windows
  • Model weighting
  • Which clients may use the seat
API

Metered API tokens

Unit
Price per 1M input, cached input, and output tokens
Examples
  • Anthropic
  • OpenAI
  • Google
  • DeepSeek
  • Mistral
  • xAI
  • hosts and gateways
What moves it
  • List price
  • Cache multipliers
  • Batch and off-peak discounts
  • Regional premiums
  • Tokenizer changes
Capacity

Owned or rented capacity

Unit
GPU-hours divided by useful work at your utilization
Examples
  • Local machines
  • rented GPUs
  • self-hosted open-weight models
What moves it
  • Hardware and rental prices
  • Model fit
  • Utilization
  • Operating staff

The FinOps loop, applied to agents

The FinOps Foundation framework phases are Inform, Optimize, and Operate.

  1. 1

    Inform

    What does a unit of accepted work cost today?

    Record tokens per attempt, cache hit rate, retries, and review time. Divide by accepted tasks, not requests.

    AI unit economics β†’
  2. 2

    Optimize

    Which lever saves the most for the least risk on this workload?

    Pick one lever, test it on a sample, compare cost per accepted task before and after.

    Lever map β†’
  3. 3

    Operate

    Who can spend what, and on which plan?

    Per-team budgets and keys, hard caps for unattended agents, seats for people, API for services.

    Subscription strategy β†’
The FinOps loop for coding agents: Inform, Optimize and Operate in a cycle around cost per accepted task, with the questions each phase answers, from what a unit of accepted work costs today to when prices are re-checked.
The full loop, with all seven questions from the guide. Read the phase-by-phase table β†’

Ten cost levers and their limits

Each lever is linked to the cost regimes it acts on. Open one to see what it leaves unsolved or can make worse, which is why each one needs a before-and-after measurement of cost per accepted task.

  1. 01 Route by task complexity APISubscription

    What it does not solve

    A task that needs the strongest model stays expensive; a wrong route costs a retry

    Acts on: API, Subscription

    Read in the guide β†’
  2. 02 Prompt caching API

    What it does not solve

    Any change in the cached prefix invalidates it; context-rewriting proxies can break it

    Acts on: API

    Read in the guide β†’
  3. 03 Batch processing API

    What it does not solve

    Processed asynchronously within a 24-hour window, so not for an interactive loop

    Acts on: API

    Read in the guide β†’
  4. 04 Off-peak pricing API Subscription (some coding plans only)

    What it does not solve

    Among providers checked, only DeepSeek (API) and Z.AI (coding plan) publish it

    Acts on: API, Subscription (some coding plans only)

    Read in the guide β†’
  5. 05 Tool output compression APISubscription

    What it does not solve

    A lossy filter can drop the line the agent needed and trigger a rerun

    Acts on: API, Subscription

    Read in the guide β†’
  6. 06 Read delegation APISubscription

    What it does not solve

    The worker model's tokens are still billed

    Acts on: API, Subscription

    Read in the guide β†’
  7. 07 Sub-agent isolation and iteration caps APISubscription

    What it does not solve

    Poorly scoped sub-agents repeat work in parallel

    Acts on: API, Subscription

    Read in the guide β†’
  8. 08 Provider portability API

    What it does not solve

    Models are not interchangeable: tool schemas and caching differ

    Acts on: API

    Read in the guide β†’
  9. 09 Plan portfolio Subscription

    What it does not solve

    Quotas stay revocable and are not guaranteed in tokens

    Acts on: Subscription

    Read in the guide β†’
  10. 10 Local or rented inference Capacity

    What it does not solve

    Only open-weight models that fit the hardware; idle hardware still costs money

    Acts on: Capacity

    Read in the guide β†’
  • Acts on this regime
  • Some coding plans only

One day of agent work, priced

The default inputs are a hypothesis, not a measured average: 20M input tokens, 80% served from cache, and 0.4M output tokens, at list prices read on September 30, 2026. The bill ranges from $0.76 to $76 for the same token count; the work each model completes is not measured here. Change the inputs to match your own usage.

Try your own numbers

  1. OpenAI GPT-6 Astra $76.00 / day $1,596.00 / month
  2. Anthropic Claude Fable 5.1 $64.00 / day $1,344.00 / month
  3. xAI Grok 4.7 $42.40 / day $890.40 / month

    Upper bound, no cached price published

  4. Anthropic Claude Opus 5.5 $27.20 / day $571.20 / month
  5. Moonshot Kimi K3 $22.80 / day $478.80 / month
  6. OpenAI GPT-6.1 Sol $15.20 / day $319.20 / month
  7. Anthropic Claude Sonnet 5.5 $15.20 / day $319.20 / month
  8. Z.AI GLM-5.3 $11.52 / day $241.92 / month
  9. Mistral Devstral 2 $8.80 / day $184.80 / month

    Upper bound, no cached price published

  10. DeepSeek deepseek-v4-pro peak $7.57 / day $158.93 / month
  11. Groq GPT-OSS 120B $2.04 / day $42.84 / month
  12. DeepSeek deepseek-flash peak $1.78 / day $37.30 / month
  13. OpenAI GPT-6 Luna $0.76 / day $15.96 / month

Prices read 2026-09-30; token counts differ between tokenizers, so the same text does not produce the same token count on every model.

This table compares bills for the same token count, not work done. Standard tier, USD per 1M tokens, cache writes excluded.

Same token count, not same outcome. A model that fails, loops on tools, or needs three attempts costs more per accepted task than its bar suggests, and tokenizers count the same text differently. Method and all rows β†’

Six claims that did not hold up

This snapshot started from an AI-generated market summary, and every claim was re-read at the vendor's own page. List prices held almost everywhere. These six did not, and they are the kind of error to expect from any AI summary of prices.

How the price snapshot was checked: claims from an AI-generated summary were re-read at vendor pricing pages, help centers, API documentation and changelogs; list prices held almost everywhere, while errors clustered in dates, plan conditions and model attribution.
  • 3 Plan condition
  • 1 Date
  • 1 Model attribution
  • 1 Not found at the source
  1. Plan condition
    The summary said

    Agent SDK and claude -p usage has its own credit since June 15, 2026

    The source says

    Announced, then paused. That usage still draws from subscription limits.

  2. Plan condition
    The summary said

    DeepSeek off-peak means nights and weekends at half price

    The source says

    Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays; everything else, weekends included, is off-peak.

  3. Plan condition
    The summary said

    OpenRouter zero data retention is a Business and Enterprise feature

    The source says

    Zero-data-retention routing is on every plan; EU in-region routing is the Business and Enterprise feature.

  4. Date
    The summary said

    Cerebras Code gives 24M tokens a day for $50

    The source says

    The figure is real, but the source post dates from November 2025 and both paid tiers were sold out on Sep 30, 2026.

  5. Model attribution
    The summary said

    Gemini 3.5 scores 76.2% on Terminal-Bench 2.1

    The source says

    The score belongs to Gemini 3.5 Flash.

  6. Not found at the source
    The summary said

    Scaleway Generative APIs start at €0.15 input, L4 GPU at €0.93 per hour

    The source says

    Not found on Scaleway's price pages, which showed different figures on Sep 30, 2026.

Go deeper

From the blog

Field reports that apply these levers to real workloads, published on the author's blog.