What usually moves the total
- Output tokens are often priced differently from input tokens.
- Heartbeats and scheduled jobs create background usage even when no chat is active.
- Fallbacks and retries can multiply a single user-visible request.
- Prompt caching changes the effective input cost only when the provider and request qualify.
- Local models can reduce API spend but add hardware, electricity, and maintenance cost.
Reconcile against actual usage
After a week, compare this estimate with provider usage exports and OpenClaw usage tracking. Update rates, tokens per run, frequency, and fallback share. Keep currency, tax, discounts, free tiers, image/audio calls, and storage explicit instead of hiding them in one total.
Common questions
Before you act
Are the default model prices current?
They are editable examples. Always confirm the linked provider page and contract before budgeting.
Does zero-cost local inference mean zero cost?
No. Hardware, power, hosting, storage, networking, and operator time still exist.
Primary sources
Verify against the owner.
Content snapshot follows official docs main at the recorded commit; the current package metadata reported 2026.8.1 when verified. Your installed release and live CLI schema remain authoritative.