Chat
Send messages to any text model with POST /chat/completions, in the OpenAI format, streamed or not.
POST /chat/completions takes the OpenAI chat completions body and returns an OpenAI chat completion. Use a model id from Models or "auto".
curl https://ewpire.com/api/v1/chat/completions \
-H "Authorization: Bearer $EWPIRE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5",
"max_tokens": 400,
"messages": [
{"role": "system", "content": "You answer in two sentences."},
{"role": "user", "content": "What is a vector database?"}
]
}'Streaming
Set "stream": true to get the answer as server-sent events: data: {chunk} lines, ending with data: [DONE]. Add "stream_options": {"include_usage": true} to get a last chunk with token usage and what the request cost.
curl -N https://ewpire.com/api/v1/chat/completions \
-H "Authorization: Bearer $EWPIRE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "auto", "stream": true, "stream_options": {"include_usage": true}, "max_tokens": 300, "messages": [{"role": "user", "content": "Write a haiku about the sea."}]}'Images in a message
Models with image in input_modalities read images. Pass them as image_url parts, as a link or a data: URL:
JSON
{
"model": "auto",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
]
}]
}With "auto", a message with an image goes to a model that can see it. A model that cannot read images answers 400 unsupported_content.
Parameters
| Parameter | Notes |
|---|---|
model | Required. A model id or "auto". |
messages | Required. 1 to 1,000 messages. Roles system, developer (read as system), user, assistant, tool. |
max_tokens / max_completion_tokens | 1 to 65,536. Default 8,192 when you send neither. |
temperature | 0 to 2. |
stream, stream_options.include_usage | See Streaming above. |
tools, tool_choice | Function tools, up to 128. |
response_format | text, json_object or json_schema. |
n | Only 1. |
top_p: 1, frequency_penalty: 0, presence_penalty: 0, logprobs: false, user and parallel_tool_calls are accepted and change nothing.
Where ewpire differs from OpenAI
- Unsupported parameters are refused, not ignored.
stop,seed,logit_bias, annabove 1, any other value of the parameters above and any unknown parameter get400with the codeunsupported_parameterand the name inparam. You never believe a setting took effect when it did not. max_tokensdefaults to 8,192. Without it an answer stops at 8,192 tokens.- What a request cost comes in headers, or in the usage chunk when streaming. See Usage.
- Size limits. The request body can be up to 8 MB (images travel inline). The prompt is limited by the model's context, not by a character count.
- Retries are safe with
Idempotency-Key. Send the header with a unique value per request; a repeat within 24 hours returns the first answer and is not charged again (it carriesIdempotent-Replayed: true). The same key with a different body gets422 idempotency_key_reused; a repeat while the first is still running gets409 idempotency_key_in_use.