A hard dollar ceiling for your OpenAI key.
Change your base URL, keep your own key, and set a limit per project. When the next request would cross it, Fuse answers 402 and the call never reaches OpenAI.
Hard spend ceiling, best-effort cost estimate. Setup takes about five minutes.
- 14:02:11chat$0.0031200 ok
- 14:02:12responses$0.0118200 ok
- 14:02:14chat$0.0027200 ok
- 14:02:15embeddings$0.0001200 ok
- 14:02:17chat—402 blown
Two environment variables. No code changes.
Fuse speaks the OpenAI API, so the official SDKs work as they are, streaming included. Your app sends a Fuse key; Fuse sends your OpenAI key.
OPENAI_BASE_URL=https://fuse.ospalabs.com/v1 OPENAI_API_KEY=fz_••••••••••••••••••••••••When the fuse blows
{
"error": "fuse_blown",
"message": "The daily ceiling for this project would be crossed by this request.",
"spent_usd": 24.97,
"limit_usd": 25,
"in_flight_usd": 0.0118,
"request_estimate_usd": 0.041,
"period": "daily",
"period_resets_at": "2026-10-11T00:00:00.000Z",
"request_id": "req_9f2c…"
}import OpenAI from "openai"; // reads OPENAI_BASE_URL and OPENAI_API_KEY const openai = new OpenAI(); try { const res = await openai.chat.completions.create({ model: "gpt-4.1-mini", messages, }); } catch (err) { if (err instanceof OpenAI.APIError && err.status === 402) { // ceiling reached: show a fallback, queue it, or tell the user } }
Every response carries x-fuse-request-id and x-fuse-reserved-usd, so you can line up a request with its row in the dashboard.
Every request is priced at its worst case before it leaves.
Four steps, in this order, on every call. If any step can't run, the request doesn't either.
Fuse counts the input with tiktoken and reads the output cap from your request: max_completion_tokens, max_tokens or max_output_tokens. No cap set? Fuse adds your project default (4,096 tokens unless you change it).
Reasoning tokens bill as output and sit inside that same cap, so they're covered too.
The worst case is reserved against the project's counter in one atomic step. If spent plus in-flight plus this request would pass the ceiling, you get a 402 and OpenAI is never called. Fifty requests arriving together can't all squeeze through the last dollar.
Your OpenAI key is decrypted in memory for the call. The request goes out unchanged apart from the output cap, and streamed tokens pass through as they arrive.
When OpenAI reports usage, the reservation is swapped for the real cost, cached input included. If a stream is cut off and usage never arrives, the full reservation counts. An error from OpenAI before any tokens are generated costs nothing.
Fails closed: a model missing from the price table gets 400 model_not_priced. If Fuse can't reach its own counter or decrypt your key, it answers 503 and forwards nothing.
What it does, and what it deliberately doesn't.
- +A hard stop per projectDaily and monthly ceilings, in your timezone, with a warning at the threshold you pick.
- +Bring your own keyYour OpenAI account, your rate limits, your OpenAI invoice. We never resell tokens.
- +Loud when it mattersAn email and a signed webhook at your warning threshold, and again on the first block.
- +Text endpoints in v1Chat completions and Responses (text, function tools), plus embeddings.
- −A prompt logNo prompt or response text is stored, in the database or in logs. No IP addresses either.
- −A firewallNo jailbreak, prompt-injection or PII filtering. It watches dollars, not words.
- −Accurate to the centCosts are estimated from reported tokens and OpenAI's public prices. Your OpenAI invoice is the final word.
- −Everything OpenAI sellsImages, audio, files, batch, realtime, built-in tools and n > 1 get a 400 and never reach OpenAI.
A flat subscription. Tokens stay between you and OpenAI.
Billed monthly through Paddle. Cancel from the billing portal whenever you like.
- Projects 3
- Fuse keys per project 3
- Requests per minute, per key 300
- Daily + monthly ceilings yes
- Email + webhook alerts yes
- Projects 20
- Fuse keys per project 10
- Requests per minute, per key 1,200
- Daily + monthly ceilings yes
- Email + webhook alerts yes