Benchmarks

We Don't Compete On Raw Proxy Speed. We Publish What Governance Costs.

Cloptima is a governance and FinOps control plane, not a pure-play LLM gateway. We're not chasing the fastest raw proxy benchmark — we publish, and keep re-publishing, what full governance actually costs in added latency, and what it guarantees when it matters more than speed.

Govern access before spend happens

Every AI gateway vendor claims to be fast. Almost none isolate governance overhead from provider variance, disclose the resource allocation behind the number, or say what happens to the number under guardrails and budget enforcement instead of a bare proxy pass-through. That makes "fast" a marketing word, not a comparable fact.

  • Vendor benchmarks usually run every policy off
  • Provider response time (which no gateway controls) gets blended into the headline number
  • Guardrails, budget enforcement, and caching are rarely measured as their own scenario
  • Resource allocation behind the number is rarely disclosed

One policy layer across model usage

Measured against a deterministic, fixed-latency backend (not a live model provider) so the number reflects Cloptima's own governance overhead, not provider variance. Gateway overhead is sourced from server-side metrics, so client network latency and provider response time are excluded from the published figure.

  • Fixed-rate load against a deterministic backend, not a live model provider
  • 30-second warmup, 3-minute measured window per rate step
  • A rate only counts as sustainable at >=99% of target traffic achieved, <=1% errors, zero dropped iterations
  • For blocked-input and model-denied, provider request counts are verified at zero before we publish
  • Only p95 is published here — p99 needs 5,000+ samples to be statistically reportable and reads noisy below that at these rates

Start with one app, then expand

Governed, guardrails, and exact-cache are performance claims: what governance and safety checks cost in added milliseconds. Blocked-input, model-denied, and budget-race are correctness claims: what governance guarantees, verified by checking that zero provider calls or zero overspend actually happened — not just that the number looked fast.

Built for private production AI

This is not a raw-proxy throughput comparison against Portkey, Bifrost, LiteLLM, or similar gateways — they're built and optimized for exactly that, and we're not trying to out-proxy the proxies. Cloptima's job is to make full governance — virtual-key auth, attribution, policy, pricing, atomic budget enforcement, guardrails, durable accounting — cost close to nothing in latency, so governance is never the reason a team skips it.

What each scenario proves

Governed, guardrails, and exact-cache are latency claims. Blocked-input, model-denied, and budget-race are correctness guarantees, verified rather than timed.

CapabilityWhat it measuresResult
Governed requestVirtual-key auth, fixed team/app/environment attribution, policy binding, live pricing, and atomic budget reservation, all before the request leaves the control plane — plus a durable, reconciled accounting record after.Publishing after our next benchmark run clears review
GuardrailsThe same governed path with input and output safety detectors enabled (prompt injection, jailbreak, toxicity, PII, secrets) against benign traffic.Publishing after our next benchmark run clears review
Exact response cacheGoverned cache hits versus misses on identical requests.Publishing after our next benchmark run clears review
Prompt-injection enforcementEnforcement decision latency for requests carrying a known prompt-injection pattern.Blocked before provider egress — zero provider calls made, on every run.
Model / provider allow-list enforcementEnforcement decision latency for a request targeting a model or provider outside the bound policy.Blocked before provider egress — zero provider calls made, on every run.
Atomic budget enforcementA concurrent burst of requests racing against the same budget ceiling at the same instant.No overspend under concurrent load — reconciled against the durable ledger, every run.

Launch path

Each scenario below states what it measures and why it matters for a governance buyer specifically, not a raw-throughput buyer. Numbers update as we re-run and re-publish the suite; this page is the single place they live.

FAQ

Operationalize LLM FinOps Across Your Apps

Start with telemetry, gateway governance, or provider bill matching workflows. Keep model spend connected to engineering ownership and finance reporting.

No credit card required
5-minute setup
Free trial