Subscription quota
- Unit
- A share of an allowance not published in tokens
- Examples
-
- Claude Pro, Max, Team
- ChatGPT Plus, Pro
- Cursor
- Copilot
- What moves it
-
- Plan multipliers
- 5-hour and weekly windows
- Model weighting
- Which clients may use the seat
In September 2026, OpenAI announced it would halve a subscription allowance while its API price list offered a model at one fifth of the Astra rate. Which cost regime carries your work decides what changed for you.
By September 30, 2026, OpenAI's API price list showed a model at one fifth of the Astra rate, and the Pro $200 allowance was set to halve. Anthropic's separate credit for programmatic usage, announced in May, is still paused. None of the three raises a per-token list price.
Each one has its own unit and moves on its own. Do not convert a quota into tokens: vendors weight allowances by model, effort, and context size.
The FinOps Foundation framework phases are Inform, Optimize, and Operate.
What does a unit of accepted work cost today?
Record tokens per attempt, cache hit rate, retries, and review time. Divide by accepted tasks, not requests.
AI unit economics βWhich lever saves the most for the least risk on this workload?
Pick one lever, test it on a sample, compare cost per accepted task before and after.
Lever map βWho can spend what, and on which plan?
Per-team budgets and keys, hard caps for unattended agents, seats for people, API for services.
Subscription strategy β
Each lever is linked to the cost regimes it acts on. Open one to see what it leaves unsolved or can make worse, which is why each one needs a before-and-after measurement of cost per accepted task.
What it does not solve
A task that needs the strongest model stays expensive; a wrong route costs a retry
Acts on: API, Subscription
Read in the guide βWhat it does not solve
Any change in the cached prefix invalidates it; context-rewriting proxies can break it
Acts on: API
Read in the guide βWhat it does not solve
Processed asynchronously within a 24-hour window, so not for an interactive loop
Acts on: API
Read in the guide βWhat it does not solve
Among providers checked, only DeepSeek (API) and Z.AI (coding plan) publish it
Acts on: API, Subscription (some coding plans only)
Read in the guide βWhat it does not solve
A lossy filter can drop the line the agent needed and trigger a rerun
Acts on: API, Subscription
Read in the guide βWhat it does not solve
The worker model's tokens are still billed
Acts on: API, Subscription
Read in the guide βWhat it does not solve
Poorly scoped sub-agents repeat work in parallel
Acts on: API, Subscription
Read in the guide βWhat it does not solve
Models are not interchangeable: tool schemas and caching differ
Acts on: API
Read in the guide βWhat it does not solve
Quotas stay revocable and are not guaranteed in tokens
Acts on: Subscription
Read in the guide βWhat it does not solve
Only open-weight models that fit the hardware; idle hardware still costs money
Acts on: Capacity
Read in the guide βThe default inputs are a hypothesis, not a measured average: 20M input tokens, 80% served from cache, and 0.4M output tokens, at list prices read on September 30, 2026. The bill ranges from $0.76 to $76 for the same token count; the work each model completes is not measured here. Change the inputs to match your own usage.
Upper bound, no cached price published
Upper bound, no cached price published
Same token count, not same outcome. A model that fails, loops on tools, or needs three attempts costs more per accepted task than its bar suggests, and tokenizers count the same text differently. Method and all rows β
This snapshot started from an AI-generated market summary, and every claim was re-read at the vendor's own page. List prices held almost everywhere. These six did not, and they are the kind of error to expect from any AI summary of prices.
Agent SDK and claude -p usage has its own credit since June 15, 2026
Announced, then paused. That usage still draws from subscription limits.
DeepSeek off-peak means nights and weekends at half price
Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays; everything else, weekends included, is off-peak.
OpenRouter zero data retention is a Business and Enterprise feature
Zero-data-retention routing is on every plan; EU in-region routing is the Business and Enterprise feature.
Cerebras Code gives 24M tokens a day for $50
The figure is real, but the source post dates from November 2025 and both paid tiers were sold out on Sep 30, 2026.
Gemini 3.5 scores 76.2% on Terminal-Bench 2.1
The score belongs to Gemini 3.5 Flash.
Scaleway Generative APIs start at β¬0.15 input, L4 GPU at β¬0.93 per hour
Not found on Scaleway's price pages, which showed different figures on Sep 30, 2026.
Field reports that apply these levers to real workloads, published on the author's blog.