Responses
POST /v1/responses is forwarded for OpenAI and xAI. Other providers return 400 endpoint_not_supported. max_output_tokens is the output cap. Text input and function tools are priced. Built-in tools are not.
The input may be a string or an array of items. Instructions are counted. Tool calls you send in the input are counted from their JSON, which overestimates on purpose. A missing max_output_tokens is filled with the project output cap before the reservation is taken.
{
"model": "gpt-4.1-mini",
"input": "Summarize the notes.",
"max_output_tokens": 400
}Anthropic, Google, and Mistral models are not available on this path. Call /v1/chat/completions for those. A responses request for a model that is off, unpriced, or missing a key fails the same way chat does, before the provider is contacted.