Windsurf quota and billing
How Windsurf bills a request depends on plan: self-serve quota (daily + weekly refresh) or enterprise ACUs. This page answers quota meter questions and the GSC fragment how-windsurf-bills-a-request.
Hub: Tokenminning in Windsurf.
Self-serve quota plans
Quota docs : allowance refreshes automatically. Consumption scales with tokens processed; cost per token varies by model.
| Model tier | Quota impact |
|---|---|
| SWE-1.5 / SWE-1.6 | Free on current quota plans |
| Claude / GPT frontier | Scales with tokens + context size |
| Fast variants | Higher per-token cost for speed |
What does not burn Cascade quota:
- Command (
Cmd/Ctrl+I) inline edits - Tab autocomplete
- Auto-generated Memories creation/retrieval (but memories still add tokens when attached)
Prompt caching
Follow-up messages in the same Cascade conversation with the same model reuse cached context at reduced cost. Switching models mid-thread loses the cache — start a new Cascade chat when changing tiers.
Where to measure
| Surface | Location |
|---|---|
| In-editor meter | Daily/weekly remaining quota |
| Plan page | windsurf.com/subscription/manage-plan |
| Per message | Cascade Stats for Nerds on chat rows |
| Teams / Enterprise | Analytics or Cascade Analytics API |
Enterprise ACUs
Enterprise may bill Agent Compute Units instead of quota. Legacy credit plans charge per Cascade message to premium models. Confirm your contract before optimizing.
Guardrails
- Enable extra-usage caps before on-demand billing
- Team admins: usage configuration API for per-user caps
- Glance at meter after heavy Cascade weeks
Related
Last updated on