Skip to content

Billing Rules

Transparent, pay-as-you-go billing. You only pay for what you use — no hidden fees, no monthly minimums.

Overview

Pay-as-you-go — no subscriptions, no monthly fees. Top up your balance and use it across any model.

Billed per token for text models, per image for image generation models. Prices listed on the Pricing page.

All amounts are calculated in micro-dollars (1 USD = 1,000,000 micros) for maximum precision.

Real-time usage tracking — view cost breakdowns by model, date, and request in your Dashboard.

How Billing Works

Every API request follows a 3-step billing flow to ensure you're never overcharged.

1
Pre-Deduct

Before calling the upstream provider, we estimate the cost based on your input tokens and temporarily hold that amount from your balance.

2
Call Upstream

Your request is sent to the AI provider. We track the actual token usage from the provider's response.

3
Settle & Refund

We compare the actual cost to the estimate. If actual < estimate, the difference is instantly refunded to your balance.

Cost Formula
cost = input_tokens × input_price + output_tokens × output_price

For text models, cost is based on input and output tokens at the rate applicable to the request. See the Pricing page for current public prices.

View Pricing
Billing Scenarios

Different outcomes are handled differently. Here's exactly what happens in each case.

Not Charged

Request fails before upstream call

Validation errors, insufficient balance, rate limit exceeded, invalid model — nothing is deducted. Full refund.
No Charge

Upstream returns an error (4xx / 5xx)

If the provider returns an error before any data is streamed, the pre-deducted amount is fully refunded.
No Charge

All providers fail (failover exhausted)

If every provider in the failover chain fails, you are not charged. Full refund of the estimated amount.
No Charge
Charged

Streaming interrupted by client

If you disconnect after streaming begins, billing uses the actual usage generated when available. If exact usage is unavailable, it is estimated from the generated content.
Charged

Streaming fails mid-stream (upstream error)

If the provider fails after streaming begins, billing uses the actual usage generated when available. If exact usage is unavailable, it is estimated from the partial response received.
Charged
Special Cases

Failover succeeds on a different provider

If the primary provider fails and a fallback succeeds, you are charged based on the fallback provider's actual token usage and pricing — not the failed attempt.
Special
Prompt Cache Pricing

Models that support prompt caching may use different cache write and cache read rates. See each model's detail page for current pricing.

Questions about billing?

Check our FAQ or view per-model pricing. Start building with transparent, predictable costs.

One API for the world’s leading AI models.

© 2026 Cloudsome. All rights reserved.