Chat completions
POST /v1/chat/completions. The body needs model and a non-empty messages array. max_tokens and max_completion_tokens are both accepted. If you set neither, Fuse inserts the project output cap and prices against that cap.
{
"model": "gemini-2.5-flash",
"messages": [{ "role": "user", "content": "hi" }],
"max_tokens": 16
}n must be 1 or omitted. Audio output, image parts, built-in web search, and predicted outputs are 400 unsupported_in_v1. Function tools you define yourself are allowed. The model name selects the provider key.
What Fuse checks, in order:
- The Fuse key, then the per-minute limit on that key.
- The account has an active plan.
- The path is one Fuse can price, and the JSON is an object with a model.
- The model is turned on, has a price, and that provider has a key.
- The worst-case cost fits under every ceiling on the project.
Only then is the provider key decrypted and the body sent. If you set stream: true, Fuse adds stream_options.include_usage on chat so the final chunk can carry token counts.
A non-OpenAI model is estimated with a padded input count, because those tokenizers are not the same as tiktoken. The reservation is the high side. The logged cost uses the usage the provider returns.