OSPA Labs / Fuse
BYOK proxy for the OpenAI API

A hard dollar ceiling for your OpenAI key.

Change your base URL, keep your own key, and set a limit per project. When the next request would cross it, Fuse answers 402 and the call never reaches OpenAI.

Hard spend ceiling, best-effort cost estimate. Setup takes about five minutes.

support-bot · dailyFuse blown
$24.97of $25.00
  1. 14:02:11chat$0.0031200 ok
  2. 14:02:12responses$0.0118200 ok
  3. 14:02:14chat$0.0027200 ok
  4. 14:02:15embeddings$0.0001200 ok
  5. 14:02:17chat—402 blown
[01] Setup

Two environment variables. No code changes.

Fuse speaks the OpenAI API, so the official SDKs work as they are, streaming included. Your app sends a Fuse key; Fuse sends your OpenAI key.

.env
OPENAI_BASE_URL=https://fuse.ospalabs.com/v1
OPENAI_API_KEY=fz_••••••••••••••••••••••••
When the fuse blows
{
  "error": "fuse_blown",
  "message": "The daily ceiling for this project would be crossed by this request.",
  "spent_usd": 24.97,
  "limit_usd": 25,
  "in_flight_usd": 0.0118,
  "request_estimate_usd": 0.041,
  "period": "daily",
  "period_resets_at": "2026-10-11T00:00:00.000Z",
  "request_id": "req_9f2c…"
}
app/api/chat/route.ts (unchanged)
import OpenAI from "openai";

// reads OPENAI_BASE_URL and OPENAI_API_KEY
const openai = new OpenAI();

try {
  const res = await openai.chat.completions.create({
    model: "gpt-4.1-mini",
    messages,
  });
} catch (err) {
  if (err instanceof OpenAI.APIError && err.status === 402) {
    // ceiling reached: show a fallback, queue it, or tell the user
  }
}

Every response carries x-fuse-request-id and x-fuse-reserved-usd, so you can line up a request with its row in the dashboard.

[02] How it decides

Every request is priced at its worst case before it leaves.

Four steps, in this order, on every call. If any step can't run, the request doesn't either.

01Price it

Fuse counts the input with tiktoken and reads the output cap from your request: max_completion_tokens, max_tokens or max_output_tokens. No cap set? Fuse adds your project default (4,096 tokens unless you change it).

Reasoning tokens bill as output and sit inside that same cap, so they're covered too.

02Reserve it

The worst case is reserved against the project's counter in one atomic step. If spent plus in-flight plus this request would pass the ceiling, you get a 402 and OpenAI is never called. Fifty requests arriving together can't all squeeze through the last dollar.

03Forward it

Your OpenAI key is decrypted in memory for the call. The request goes out unchanged apart from the output cap, and streamed tokens pass through as they arrive.

04Settle it

When OpenAI reports usage, the reservation is swapped for the real cost, cached input included. If a stream is cut off and usage never arrives, the full reservation counts. An error from OpenAI before any tokens are generated costs nothing.

Fails closed: a model missing from the price table gets 400 model_not_priced. If Fuse can't reach its own counter or decrypt your key, it answers 503 and forwards nothing.

[03] Scope

What it does, and what it deliberately doesn't.

Fuse is
  • +
    A hard stop per projectDaily and monthly ceilings, in your timezone, with a warning at the threshold you pick.
  • +
    Bring your own keyYour OpenAI account, your rate limits, your OpenAI invoice. We never resell tokens.
  • +
    Loud when it mattersAn email and a signed webhook at your warning threshold, and again on the first block.
  • +
    Text endpoints in v1Chat completions and Responses (text, function tools), plus embeddings.
Fuse isn't
  • −
    A prompt logNo prompt or response text is stored, in the database or in logs. No IP addresses either.
  • −
    A firewallNo jailbreak, prompt-injection or PII filtering. It watches dollars, not words.
  • −
    Accurate to the centCosts are estimated from reported tokens and OpenAI's public prices. Your OpenAI invoice is the final word.
  • −
    Everything OpenAI sellsImages, audio, files, batch, realtime, built-in tools and n > 1 get a 400 and never reach OpenAI.
[04] Pricing

A flat subscription. Tokens stay between you and OpenAI.

Billed monthly through Paddle. Cancel from the billing portal whenever you like.

Starter
$19/ month
  • Projects 3
  • Fuse keys per project 3
  • Requests per minute, per key 300
  • Daily + monthly ceilings yes
  • Email + webhook alerts yes
Start with Starter
Team
$49/ month
  • Projects 20
  • Fuse keys per project 10
  • Requests per minute, per key 1,200
  • Daily + monthly ceilings yes
  • Email + webhook alerts yes
Start with Team
[05] Questions

Things people ask before switching the URL.

Can spend go past the ceiling?
Every request reserves its worst case before it's sent, so in normal operation committed spend stays under the line. Two things can nudge it over: OpenAI counting a few more tokens than tiktoken did for the same text, and a price change we haven't picked up yet. Both are small, which is why we call it a hard ceiling with a best-effort estimate rather than promising more.
Does Fuse see my prompts?
It has to pass them to OpenAI, so they sit in memory for the length of the request. They are never written to a database, a log file or an analytics tool. The dashboard shows metadata only: time, endpoint, model, token counts, cost and status.
How is my OpenAI key stored?
Encrypted with AES-256-GCM under its own data key, which is wrapped by a master key that lives only in Cloudflare Secrets. The dashboard shows the last four characters. You can replace or revoke a key; nobody can view it again, including us.
What does it add to latency?
One extra network hop through Cloudflare and a check against your project's counter. Streams are passed through chunk by chunk, not buffered.
What happens if Fuse itself has a problem?
It fails closed. If the counter, the price table or key decryption is unavailable, Fuse answers 503 and sends nothing to OpenAI. A fuse that lets traffic through when it's unsure isn't doing its job.
Anthropic, team seats, per-route limits?
On the list for v2. v1 is OpenAI only, one owner per account.