Notes from the engineers who actually run the gateway, written against the codebase rather than the roadmap. Where an AI bill really leaks, what happens inside the router in twenty milliseconds, and what provable zero-retention costs to build.
Long-form pieces written against the codebase. Filter by topic, or read them in order — they build on each other.
"We don't store your data" is a claim. A per-request retention header, a customer-held encryption key, and a crypto-shred certificate are proof.
AI spend rarely explodes in one moment. It leaks — on every call, every key, every day — until the model line is 4× last quarter and nobody can say why.
A committed API key or a compromised service account can burn six or seven figures in a single day. Models are fast enough to spend faster than any human notices.
A one-line classification and a paragraph of literary analysis are not the same task. Most gateways price them the same anyway.
Token compression is the cost lever nobody remembers to pull — right up until a debug session with a 40,000-line log attached shows up on the invoice.
Both sit between your app and every model provider. Where they actually differ — routing depth, retention posture, self-hosting, and how each prices — matters more than either company's marketing.