Skip to content

API keys are not open yet. These docs describe the API as it will work when keys open, so you can plan your integration now.

Docs

Usage

What each request costs, how it is charged to your balance, and how to read your usage.

How a request is charged

Every request is paid in credits from the balance of the key's workspace, at the model's price per million input and output tokens (see Models and Pricing). Before a request runs, ewpire reserves its most it could cost: the prompt plus max_tokens at the model's price, or at the dearest model Auto might use. When the answer is done, only what was actually used is charged and the rest of the reservation is released.

If your balance cannot cover the reservation, the request is refused with 429 insufficient_quota and nothing is charged. With a small balance, a lower max_tokens makes the reservation smaller.

A failed request is not charged. A stream you stop after the answer has started is charged for the tokens already produced.

What a request cost

A non-streamed answer carries two headers, both in credits with six decimals:

HeaderMeaning
X-Ewpire-Credits-ChargedWhat this request cost.
X-Ewpire-BalanceCredits available after it.
X-Request-IdThe id of the request in your usage records.

A streamed answer cannot change its headers once it has started, so the cost comes in the last chunk instead, when you ask for "stream_options": {"include_usage": true}:

JSON

{
  "id": "chatcmpl-…",
  "object": "chat.completion.chunk",
  "choices": [],
  "usage": { "prompt_tokens": 14, "completion_tokens": 22, "total_tokens": 36 },
  "ewpire": { "credits_charged": "0.000262", "balance": "4.998510" }
}

Reading your usage

GET /usage returns the API usage of the key's workspace: totals, one entry per day and one per model. Pass range=7, 30 or 90 (days, the default is 30), or from and to as YYYY-MM-DD (UTC, both included, at most a year apart).

curl "https://ewpire.com/api/v1/usage?range=7" \
  -H "Authorization: Bearer $EWPIRE_API_KEY"

One request is one call to /chat/completions: refusals count, and an Auto request that moved to a second model counts once. The same numbers, with charts and a filter per key, are under API in the app.