Skip to content

API keys are not open yet. These docs describe the API as it will work when keys open, so you can plan your integration now.

Docs

Chat

Send messages to any text model with POST /chat/completions, in the OpenAI format, streamed or not.

POST /chat/completions takes the OpenAI chat completions body and returns an OpenAI chat completion. Use a model id from Models or "auto".

curl https://ewpire.com/api/v1/chat/completions \
  -H "Authorization: Bearer $EWPIRE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "max_tokens": 400,
    "messages": [
      {"role": "system", "content": "You answer in two sentences."},
      {"role": "user", "content": "What is a vector database?"}
    ]
  }'

Streaming

Set "stream": true to get the answer as server-sent events: data: {chunk} lines, ending with data: [DONE]. Add "stream_options": {"include_usage": true} to get a last chunk with token usage and what the request cost.

curl -N https://ewpire.com/api/v1/chat/completions \
  -H "Authorization: Bearer $EWPIRE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "auto", "stream": true, "stream_options": {"include_usage": true}, "max_tokens": 300, "messages": [{"role": "user", "content": "Write a haiku about the sea."}]}'

Images in a message

Models with image in input_modalities read images. Pass them as image_url parts, as a link or a data: URL:

JSON

{
  "model": "auto",
  "messages": [{
    "role": "user",
    "content": [
      {"type": "text", "text": "What is in this picture?"},
      {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
    ]
  }]
}

With "auto", a message with an image goes to a model that can see it. A model that cannot read images answers 400 unsupported_content.

Parameters

ParameterNotes
modelRequired. A model id or "auto".
messagesRequired. 1 to 1,000 messages. Roles system, developer (read as system), user, assistant, tool.
max_tokens / max_completion_tokens1 to 65,536. Default 8,192 when you send neither.
temperature0 to 2.
stream, stream_options.include_usageSee Streaming above.
tools, tool_choiceFunction tools, up to 128.
response_formattext, json_object or json_schema.
nOnly 1.

top_p: 1, frequency_penalty: 0, presence_penalty: 0, logprobs: false, user and parallel_tool_calls are accepted and change nothing.

Where ewpire differs from OpenAI

  • Unsupported parameters are refused, not ignored. stop, seed, logit_bias, an n above 1, any other value of the parameters above and any unknown parameter get 400 with the code unsupported_parameter and the name in param. You never believe a setting took effect when it did not.
  • max_tokens defaults to 8,192. Without it an answer stops at 8,192 tokens.
  • What a request cost comes in headers, or in the usage chunk when streaming. See Usage.
  • Size limits. The request body can be up to 8 MB (images travel inline). The prompt is limited by the model's context, not by a character count.
  • Retries are safe with Idempotency-Key. Send the header with a unique value per request; a repeat within 24 hours returns the first answer and is not charged again (it carries Idempotent-Replayed: true). The same key with a different body gets 422 idempotency_key_reused; a repeat while the first is still running gets 409 idempotency_key_in_use.