Chat completions

POST /v1/chat/completions. The body needs model and a non-empty messages array. max_tokens and max_completion_tokens are both accepted. If you set neither, Fuse inserts the project output cap and prices against that cap.

{
  "model": "gemini-2.5-flash",
  "messages": [{ "role": "user", "content": "hi" }],
  "max_tokens": 16
}

n must be 1 or omitted. Audio output, image parts, built-in web search, and predicted outputs are 400 unsupported_in_v1. Function tools you define yourself are allowed. The model name selects the provider key.

What Fuse checks, in order:

  1. The Fuse key, then the per-minute limit on that key.
  2. The account has an active plan.
  3. The path is one Fuse can price, and the JSON is an object with a model.
  4. The model is turned on, has a price, and that provider has a key.
  5. The worst-case cost fits under every ceiling on the project.

Only then is the provider key decrypted and the body sent. If you set stream: true, Fuse adds stream_options.include_usage on chat so the final chunk can carry token counts.

A non-OpenAI model is estimated with a padded input count, because those tokenizers are not the same as tiktoken. The reservation is the high side. The logged cost uses the usage the provider returns.