Billing Rules
Transparent, pay-as-you-go billing. You only pay for what you use — no hidden fees, no monthly minimums.
Overview
Pay-as-you-go — no subscriptions, no monthly fees. Top up your balance and use it across any model.
Billed per token for text models, per image for image generation models. Prices listed on the Pricing page.
All amounts are calculated in micro-dollars (1 USD = 1,000,000 micros) for maximum precision.
Real-time usage tracking — view cost breakdowns by model, date, and request in your Dashboard.
How Billing Works
Every API request follows a 3-step billing flow to ensure you're never overcharged.
Pre-Deduct
Before calling the upstream provider, we estimate the cost based on your input tokens and temporarily hold that amount from your balance.
Call Upstream
Your request is sent to the AI provider. We track the actual token usage from the provider's response.
Settle & Refund
We compare the actual cost to the estimate. If actual < estimate, the difference is instantly refunded to your balance.
Cost Formula
For text models, cost is based on input and output tokens at the rate applicable to the request. See the Pricing page for current public prices.
View PricingBilling Scenarios
Different outcomes are handled differently. Here's exactly what happens in each case.
Request fails before upstream call
Validation errors, insufficient balance, rate limit exceeded, invalid model — nothing is deducted. Full refund.Upstream returns an error (4xx / 5xx)
If the provider returns an error before any data is streamed, the pre-deducted amount is fully refunded.All providers fail (failover exhausted)
If every provider in the failover chain fails, you are not charged. Full refund of the estimated amount.Streaming interrupted by client
If you disconnect after streaming begins, billing uses the actual usage generated when available. If exact usage is unavailable, it is estimated from the generated content.Streaming fails mid-stream (upstream error)
If the provider fails after streaming begins, billing uses the actual usage generated when available. If exact usage is unavailable, it is estimated from the partial response received.Failover succeeds on a different provider
If the primary provider fails and a fallback succeeds, you are charged based on the fallback provider's actual token usage and pricing — not the failed attempt.Prompt Cache Pricing
Models that support prompt caching may use different cache write and cache read rates. See each model's detail page for current pricing.
Questions about billing?
Check our FAQ or view per-model pricing. Start building with transparent, predictable costs.