๐Ÿ›ฐ๏ธ

AI Agent Routing / Gateway Cost Comparator

Compare AI gateway and LLM routing costs across leading platforms. Model gateway fees, LLM token spend, caching savings, routing savings, and margin pricing.

LLM workload

Gateway platform

Core gateway is free. Persistent logs limited (100k free, 10M paid per gateway). Logpush billed at $0.05/million on Workers Paid. Provider token costs pass through.

Counted only for self-hosted option.

Pricing model

Total monthly cost

โ€”

Net LLM spend / mo

โ€”

Gateway fee / mo

โ€”

12-month total

โ€”

Raw LLM spend / mo

โ€”

Cache savings / mo

โ€”

Routing savings / mo

โ€”

Self-host compute / mo

โ€”

Engineering / mo

โ€”

Pricing to hit margin

Required monthly revenue

โ€”

Price per 1k requests

โ€”

Price per 1M tokens

โ€”

Cost comparison by gateway provider

Same workload and LLM spend across all gateway options.

Gateway / plan Base Overage Logpush Compute Eng. Total/mo
Cloudflare AI Gateway
Core gateway is free. Persistent logs limited (100k free, 10M paid per gateway). Logpush billed at $0.05/million on Workers Paid. Provider token costs pass through.
โ€” โ€” โ€” โ€” โ€” โ€”
Portkey AI Gateway
$49/mo base with 100k recorded logs. +$9 per 100k logs up to 3M. OSS self-hosted available. Provider token costs pass through.
โ€” โ€” โ€” โ€” โ€” โ€”
Helicone AI Gateway
$79/mo Pro with ~1M requests. Free tier for 10k/mo. Zero markup on provider tokens. Apache-2.0 self-host option.
โ€” โ€” โ€” โ€” โ€” โ€”
LiteLLM Enterprise (managed)
Enterprise Basic starts around $250/mo. Premium around $2,500/mo. Pricing by annual request volume and support level. OSS core is free self-hosted.
โ€” โ€” โ€” โ€” โ€” โ€”
Braintrust Gateway
Gateway currently free in beta. Starter plan free with 1M spans + 1 GB data; Pro $249/mo. Tracing and eval platform separate.
โ€” โ€” โ€” โ€” โ€” โ€”
Zuplo (API gateway for AI)
$0 free tier with 100k requests/mo; Builder plan $25/mo with scalable limits. Token-based rate limits and policies.
โ€” โ€” โ€” โ€” โ€” โ€”
TrueFoundry AI Gateway
$499/mo Pro includes 1M requests. Enterprise custom. MCP gateway, governance, and routing included.
โ€” โ€” โ€” โ€” โ€” โ€”
Kong AI Gateway (Konnect)
~$105/service/mo plus usage; AI Gateway add-on; model proxy $100/model/mo in Konnect. Enterprise custom. Best for existing Kong users.
โ€” โ€” โ€” โ€” โ€” โ€”
Self-hosted LiteLLM / Bifrost OSS
Open-source self-hosted: no per-request gateway fees. You pay compute, storage, and engineering time. Enterprise license optional.
โ€” โ€” โ€” โ€” โ€” โ€”

How the estimate works

  • Raw LLM spend = requests ร— ((input tokens รท 1M ร— input price) + (output tokens รท 1M ร— output price)).
  • Cache savings = raw LLM spend ร— cache hit% ร— cache cost reduction%.
  • Routing savings = raw LLM spend ร— routing savings%.
  • Net LLM spend = raw LLM spend โˆ’ cache savings โˆ’ routing savings.
  • Gateway fee = base fee + request overage + logpush. Many tiers have zero gateway fee.
  • Margin pricing = total monthly cost รท (1 โˆ’ margin%).

When to use each gateway

  • โ€ข Cloudflare for free core gateway if already on Cloudflare Workers/Pages.
  • โ€ข Portkey when you need managed gateway + prompt management + guardrails.
  • โ€ข Helicone for open-source-friendly gateway with zero provider markup.
  • โ€ข LiteLLM / Bifrost OSS if you want self-hosting and multi-provider standardization.
  • โ€ข Braintrust when gateway traffic should feed evals and release checks.
  • โ€ข TrueFoundry / Kong for enterprise governance, MCP, or existing API-management stacks.

Frequently asked questions

What does an AI gateway actually save?โ–ผ

Three main levers: (1) caching reduces repeated identical calls, (2) routing picks cheaper or faster providers for the same model class, and (3) rate limits and retries stop runaway spend from loops or bad deploys.

Do gateway fees replace LLM provider fees?โ–ผ

No for most gateways. Cloudflare, Helicone, Portkey, and Braintrust pass through provider token costs with zero markup; you pay the gateway/platform fee on top. LiteLLM OSS and Bifrost OSS have no per-token gateway fee beyond your own infrastructure.

When is a self-hosted gateway cheaper than managed?โ–ผ

At high request volume, managed log/request fees and seat costs dominate. Self-hosting LiteLLM or Bifrost shifts cost to compute and SRE time. Break-even often appears above a few million requests per month, especially with heavy caching.

How do I model caching savings?โ–ผ

Set a cache hit percentage and a cost-reduction percentage. A 20% cache hit with 80% cost reduction means 16% of otherwise-billable LLM spend is avoided. The tool subtracts that from the raw LLM estimate and shows net LLM spend.

Which gateway should I start with?โ–ผ

Cloudflare AI Gateway is free and generous for apps already on Cloudflare. Helicone and Portkey are strong managed options with observability. LiteLLM/Bifrost are good if you need self-hosting or multi-provider standardization.

What about routing savings?โ–ผ

Smart routing can move traffic to cheaper provider keys (e.g. AWS Bedrock vs direct Anthropic), use smaller models for simple prompts, or fall back to cached responses. Savings are typically 5โ€“20% of LLM spend.

Prices are directional estimates based on public list pricing for AI gateway / LLM routing platforms. Confirm current rates before budgeting. LLM provider token costs are passed through by most gateways; this tool focuses on gateway fees, infrastructure, and savings from caching/routing. Some vendors quote custom Enterprise pricing. Last updated: 2026-08-10.

๐Ÿš€ Get AI automation insights daily

15:00 MST. One-click unsubscribe.

Subscribe