Self-hosted and air-gapped.

The AI Gateway
for platform teams

Put your full AI stack behind one key. See who is driving spend, cap it before it runs, and send each request to the model that should handle it.

Self-host in minutes. No credit card.

Run in production by the teams shipping AI at scale.
Testimonial
Testimonial
Testimonial

“LiteLLM gives NVIDIA engineers a single, consistent way to access more than 100 AI model endpoints.”

Ajay Dogra
Ajay DograProduct, NVIDIA
Testimonial
Testimonial

“LiteLLM streamlines the complexities of managing multiple LLM models.”

Mark Koltnuk
Mark KoltnukPrincipal Architect, Lemonade
Testimonial
Testimonial
Testimonial

“LiteLLM has let my team provide the latest LLM models to our users, usually within a day of them being released… it has saved us months of work.”

David Leen
David LeenStaff Software Engineer, Netflix
Testimonial
Testimonial

“If we decide to switch the backend model, it’s a simple configuration update in the gateway; no code changes, procurement cycles, or repetitive security reviews required.”

Dennis Henry
Dennis HenryProductivity Architect, Okta
Testimonial
LiteLLM abstract hero background
Why LiteLLM

Give your whole org access to every model, agent, and MCP, then see, control, and optimize every request.

01

Access: Give your whole company every model.

Put every model, agent, and MCP behind one API and one login. We handle key management and work with the secret manager you already run, so platform teams open access to the whole org without becoming the bottleneck, and developers build in minutes, not a procurement cycle.

  • One OpenAI-compatible API to 140+ providers and 1,800+ models. Swap models without changing app code.
  • One login with SSO, scoped by team, project, or app.
  • Day-zero support for new models, so developers get the latest the day it ships.
  • Reach agents and MCP servers through the same gateway, not just LLMs.
  • Bring your own internal, fine-tuned, and self-hosted models behind the same key.
Virtual Keys Ask AI Export
+ Create New KeyShowing 1–6 of 6Filters
KeyTeamLast ActiveSpend / Budget
acme-prod-gatewayActivesk-…9f2Aplatform-eng2 min ago$4,182.55 / $10,000
wayne-data-scienceActivesk-…Tt09data-science3 min ago$6,740.15 / $12,000
hooli-data-squadActivesk-…Pq77ai-gatewayjust now$3,301.42 / $8,000
globex-genai-platformActivesk-…7b01ml-platform5 min ago$1,905.20 / $5,000
ci-eval-runnerRate-limitedsk-…Hh55infra7 min ago$58.44 / $250
umbrella-eng-sandboxExpiredsk-…WW21security38 min ago$121.87 / $500
Rows per page: 6Page 1 of 1
Key Ownership
Owned By
YouService AccountAnother UserAgent New
OrganizationAll Organizations
TeamSearch or select a team
Key Details
*Key Namee.g. prod-gateway-key
ModelsSelect models
Leave empty to allow access to all models
Key TypeAI APIs
Optional Settings
CancelCreate Key
Team: All TeamsView: All8 deployments
Public NameProviderModeInput / OutputHealth
gpt-5openai/gpt-5OpenAIChat$1.25 / $10Healthy
gpt-5-miniopenai/gpt-5-miniOpenAIChat$0.25 / $2Healthy
claude-opus-4-8anthropic/claude-opus-4-8AnthropicChat$5 / $25Healthy
claude-sonnet-4-5anthropic/claude-sonnet-4-5AnthropicChat$3 / $15Healthy
gemini-2.5-provertex_ai/gemini-2.5-proVertex AIChat$1.25 / $10Healthy
llama-4-maverickbedrock/llama-4-maverickBedrockChat$0.35 / $1.40Healthy
text-embedding-3-largeopenai/text-embedding-3-largeOpenAIEmbedding$0.13 / —Healthy
claude-opus-4-8 (global)anthropic/claude-opus-4-1AnthropicChat$5 / $25Degraded
Group by: Provider6 providers · 15 models
Anthropic — Bedrockclaude-opus-4-8, claude-sonnet-4-5, claude-haiku-4-5bedrock3
Anthropic — Globalclaude-opus-4-8, claude-opus-4-1-20250805global2
OpenAIgpt-5, gpt-5-mini, gpt-5.5-pro, text-embedding-3-largedirect4
Vertex AIgemini-2.5-pro, gemini-2.5-flashgoogle2
Bedrock / Cloudflarellama-4-maverick, mistral-largerouter2
RRoutersprod-router (load-balanced), cheap-fallbackfallback2
+ Create TeamShowing 1–7 of 7Org: All
TeamMembersModelsKeysSpend / Budget
Pplatform-engteam-2f9a…12All8$18.3K / $40K
Ddata-scienceteam-71c4…965$6.7K / $12K
Aai-gatewayteam-08be…6All4$3.3K / $8K
Mml-platformteam-5d20…846$1.9K / $5K
Ssupport-aiteam-9a11…433$2.6K / ∞
Ssecurityteam-c7f3…522$0.1K / $0.5K
Rresearchteam-3b8d…3All2$0.4K / $2K
Pplatform-engteam-2f9a03c1…ActiveEdit Team
Spend
$18,309
Budget
$40,000
Members
12
Keys
8
ModelsAll Proxy Models
Rate LimitsTPM 2,000,000·RPM 10,000
Guardrailspii-maskingprompt-injectionsecret-detection
Object PermissionsVector Stores · 2MCP Servers · 3
Organizationacme-inc
+ Add Member7 membersTeam: platform-eng
MemberTeam RoleLast Active
ARAna Riveraa.rivera@acme.ioAdmin2 min ago
JMJae Moraj.mora@acme.ioMember18 min ago
SPSam Patels.patel@acme.ioMember1 hr ago
KSKir Singhk.singh@acme.ioMember3 hr ago
DVDevOps Botdevops@acme.ioAdminjust now
CIci-eval-runnerservice accountService7 min ago
LCLee Chenl.chen@acme.ioMember5 hr ago
Team: AllAccess Group: All6 servers
github-mcpSSE
https://mcp.acme.io/github
14 toolsplatform
jira-mcpSSE
https://mcp.acme.io/jira
9 toolsproduct
confluence-mcpHTTP
https://mcp.acme.io/confluence
6 toolsproduct
gitlab-mcpSSE
https://mcp.acme.io/gitlab
11 toolsplatform
Ssentry-mcpHTTP
https://mcp.acme.io/sentry
7 toolssre
Sslack-mcpSSE
https://mcp.acme.io/slack
5 toolsgeneral
ModelsMCPSkillsProvider: All
ModelProviderModeContextIn / Out
gpt-5OpenAIChat400K$1.25 / $10
claude-opus-4-8AnthropicChat200K$5 / $25
claude-sonnet-4-5AnthropicChat200K$3 / $15
gemini-2.5-proVertex AIChat1M$1.25 / $10
llama-4-maverickBedrockChat128K$0.35 / $1.40
mistral-largeCloudflareChat128K$2 / $6
text-embedding-3-largeOpenAIEmbedding8K$0.13 / —
02
Key — Guardrails & Limits Ask AI Export
Rate Limits & Budgets
TPM Limit2,000,000
RPM Limit10,000
Max Budget$10,000 / month
Per-tag budgets: prod $6K · staging $2K · eval $500
Guardrails & Policies
Guardrails
pii-maskingprompt-injectionsecret-detection
Access Groupsprod-gateway, internal-tools
Pass-through RoutesAdd /v1/messages, /v1/embeddings
Advanced Policies · logging, fallbacks, model aliases
CancelSave Settings
Range: Last 7 daysTeam: All Teams1.24M requests
Requests
1.24M
Successful
99.2%
Tokens
3.8B
Spend
$33.9K
ModelRequestsTokensSpend
gpt-5openai/gpt-5512K1.6B$14.2K
claude-opus-4-8anthropic/claude-opus-4-8208K820M$9.8K
claude-sonnet-4-5anthropic/claude-sonnet-4-5301K940M$5.1K
gemini-2.5-provertex_ai/gemini-2.5-pro142K380M$3.2K
llama-4-maverickbedrock/llama-4-maverick77K210M$1.6K
+ Add Guardrail6 guardrailsMode: All
GuardrailProvider / TypeModeStatus
Ppresidio-piigr-1a4f…Presidio PIIpre-callActive
Llakera-guardgr-77b2…Lakera prompt-injectionduring-callActive
Aaporia-policygr-0c31…Aporia moderationpost-callActive
Ssecret-detectorgr-9d80…Regex secret-detectionpre-callActive
Ttoxicity-filtergr-4e12…LLM Guard toxicitypost-callDisabled
prompt-shieldgr-2f56…Bedrock Guardrailspre-callActive
Action: AllActor: All7 events today
TimestampActorActionObjectChanged By
14:32:07a.rivera@acme.iocreatedkey/acme-prod-gatewayAna Rivera
14:28:51j.mora@acme.ioupdatedteam/data-scienceJae Mora
14:19:03devops@acme.iorotatedkey/ci-eval-runnerDevOps Bot
13:58:44s.patel@acme.iodeletedguardrail/toxicity-filterSam Patel
13:47:12a.rivera@acme.ioupdatedbudget/platform-engAna Rivera
13:21:39k.singh@acme.iocreatedkey/globex-genaiKir Singh
12:55:02devops@acme.ioupdatedmcp/github-mcpDevOps Bot
Range: Last 24h84.2K evaluations
Evaluations
84.2K
Blocked
1,204
Pass Rate
98.6%
Avg Latency
42ms
Active
5 / 6
GuardrailEvaluationsBlockedPass Rate
Ppresidio-pii38.1K40298.9%
Llakera-guard21.5K61197.2%
Aaporia-policy12.9K9899.2%
Ssecret-detector7.4K7199.0%
prompt-shield4.3K2299.5%
LiveStatus: All8 most recent
TimeModelKeyStatusLatencyTokensCost
14:32:41gpt-5acme-prod200812ms2,104$0.03
14:32:39claude-opus-4-8wayne-ds2001.2s3,880$0.08
14:32:37claude-sonnet-4-5hooli-data200640ms1,510$0.02
14:32:35gemini-2.5-proglobex-genai42944ms$0.00
14:32:33gpt-5ci-eval200902ms2,340$0.03
14:32:31llama-4-maverickml-platform200388ms980$0.01
14:32:29mistral-largesupport-ai5005.0s$0.00
14:32:27claude-sonnet-4-5research200705ms1,220$0.02

Visibility & Control: See every request, and cap it before it runs.

See who and what is driving usage and spend, attribute every request for chargeback, and cap budgets before they run. Set budgets and rate limits per team; when they hit the cap, requests stop.

  • Usage and spend tracked per key, user, team, org, tool, agent, and MCP, across 140+ providers.
  • Enterprise chargeback: attribute every request and bill teams and business units for what they use.
  • Hard budgets per key, team, org, and model, with daily and monthly resets. At the cap, requests stop.
  • Rate limits and leaked-key protection, so no runaway job or compromised key runs up your bill.
  • Model access control and guardrails, with an audit log on every request.
03

Cost optimization: Maximize the ROI of your AI.

Swap models without changing a line of code, and let Auto Routing send each request to the model that should handle it, so the budget you set goes further.

  • Load balancing across providers, regions, and keys.
  • Lowest-cost routing to the cheapest deployment that can serve the request.
  • Auto Routing that sends simple prompts to cheaper models and hard ones to stronger models.
  • Response and semantic caching (Redis, S3, GCS), so you never pay twice for the same answer.
  • Prompt compression, so you send fewer tokens for the same result.
Usage — Cost by Team, Key & Model Ask AI Export
Range: Last 30 daysOrg: All$48,210.94 total
Spend by Team
platform-eng$18,309.44
data-science$10,742.15
ai-gateway$6,301.42
ml-platform$5,905.20
Top Virtual Keys
wayne-data-science$6,740.15
acme-prod-gateway$4,182.55
hooli-data-squad$3,301.42
globex-genai$1,905.20
Top Models
gpt-5$16,884.20
claude-opus-4-8$12,402.55
claude-sonnet-4-5$8,115.90
gemini-2.5-pro$5,340.18
Budget
Max Budget$10,000.00
Reset Budgetmonthly
Resets on the 1st of each month at 00:00 UTC
Budget Windowrolling 30d
Soft Limit Alert80% · #ai-platform-ops
Fallbacks & Enforcement
Budget Fallbacks
gpt-5-miniclaude-haiku-4-5
On Budget Exceeded
ThrottleBlockFallback
Advanced · per-model budgets, tag budgets, reset webhooks
CancelSave Settings
Range: Year to date7 teams
Total Spend
$48,210.94
Requests
2,847,103
Successful
2,841,569
Failed
5,534
Tokens
5,912,441,077
TeamRequestsTokensSpendBudget
Pplatform-eng1,084,2202.25B$18,309.44$40,000
Ddata-science632,9051.32B$10,742.15$18,000
Aai-gateway371,118770M$6,301.42$10,000
Mml-platform348,004720M$5,905.20$8,000
Ssupport-ai213,441440M$3,618.77$6,000
Rresearch126,510260M$2,152.09$4,000
Ssecurity70,905150M$1,181.87$2,000
PersonalOrganizationTeamTagTeam: All Teams
Scope
Team
Total Spend
$48,210.94
Requests
2,847,103
Tags
6
Spend by Tag
TagApplied ToRequestsSpend
env:prodall teams1,842,004$31,402.18
env:stagingplatform-eng, infra401,552$7,118.44
service:chatsupport-ai, ai-gateway288,110$4,290.05
service:embeddata-science171,904$2,641.60
env:devall teams96,220$1,802.55
service:agentsml-platform, research47,313$956.12
Tag budgets enforced at request time6 of 6 tags
Project SpendJul 8 – Jul 15, 2026Range: Last 8 days
Total Spend
$48,210.94
Max Budget
$250,000.00
0809101112131415
Daily spend · all teams, all modelsPeak $8,120.44 on Jul 11
04

Ease of deployment: Go live in your stack in an afternoon

Self-host the same open-source gateway behind 240M+ Docker pulls, in your own cloud or fully air-gapped. Simple enough that it just works.

  • Deploy with an official Helm chart or Terraform module
  • Official Docker images, including database-bundled and non-root variants
  • Runs on your own Postgres and Redis, and scales out with Kubernetes autoscaling
  • Self-host anywhere, including fully air-gapped environments
  • One-click deploy into your hyperscaler (AWS, GCP, or Azure)
deploy — bashCopy
$ docker pull docker.litellm.ai/berriai/litellm:latest
$ docker run -p 4000:4000 litellm
05

Sub-millisecond overhead. Read the benchmark.

The LiteLLM Rust AI Gateway is live. On our benchmarks it adds 0.66 ms at p99 — 3.5× lower overhead than the next AI gateway — measured with AI Gateway Bench, an open standard for benchmarking and comparing AI gateways. Run it yourself.

AI gateway benchmark results measured with AI Gateway Bench. LiteLLM (Rust) adds 0.66 ms at p99. Throughput 2,800+ req/s (at ~21% CPU); ~4.5× more requests per dollar. Identical hardware; every gateway pointed at the same deterministic mock upstream; single client.
AI gatewayp99 added latencyMemory at restThroughputCPU utilization
LiteLLM (Rust)0.66 ms~22 MB2,800+ req/s~21%
Portkey2.29 ms
Bifrost4.54 ms~199 MB
0.66 ms
p99 added latency
~22 MB
memory at rest
2,800+ req/s
at ~21% CPU
~4.5×
more requests per dollar
Overhead the gateway adds — p99 ms (lower is better)
LiteLLM (Rust)0.66 ms
Portkey2.29 ms
Bifrost4.54 ms

Identical hardware; every gateway pointed at the same deterministic mock upstream; single client.

Memory at rest — peak RSS (lower is better)
LiteLLM (Rust)~22 MB
Bifrost~199 MB

Same benchmark run; resident memory with the gateway idle.

Measured on identical hardware, with every gateway pointed at the same deterministic upstream, so only the gateway’s own overhead is left. Every number, chart, and script is public.

See the benchmark
140+
LLM providers
1,892
Unique models
240M+
Docker pulls
1B+
Requests served
53K+
GitHub stars
1,005+
GitHub contributors
Running in production
Single access source with consistency
LiteLLM gives NVIDIA engineers a single, consistent way to access more than 100 AI model endpoints across cloud providers, open source deployments, and internal NVIDIA services.
Ajay Dogra
AD
Ajay Dogra
Product
New models in a day, not hours of rework
LiteLLM has let my team provide the latest LLM models to our users, usually within a day of them being released… it has saved us months of work.
David Leen
David Leen
Staff Software Engineer
Multiple models, one interface
LiteLLM streamlines the complexities of managing multiple LLM models.
Mark Koltnuk
Mark Koltnuk
Principal Architect, GenAI Platform
Swap models without new security reviews
If we decide to switch the backend model, it’s a simple configuration update in the gateway; no code changes, procurement cycles, or repetitive security reviews required.
Dennis Henry
Dennis Henry
Productivity Architect
06

Learn more about our plans.

$0
Free forever.
Includes:
  • 140+ LLM provider integrations
  • Langfuse, Arize Phoenix, LangSmith, and OTEL logging
  • Virtual keys, budgets, and teams
  • Load balancing and RPM/TPM limits
  • LLM guardrails
No credit card. Self-host in minutes.
Get started
Get more with Enterprise
For teams rolling out LLM access across many developers and projects.
  • Everything in open source
  • Enterprise support and custom SLAs
  • JWT auth, SSO, and audit logs
  • All enterprise features. See the docs.
07

Your keys, your infra, your audit trail.

Security fixes, stable releases, and what we're working on next. All public. Don't take our word for it; read the changelog and the code.

  • Every release, in the open — see exactly what changed, in the public changelog.
  • A stable release track — production images ship after load testing.
  • Security in the open — when there's an issue, we publish the disclosure and the fix. You can read every one.
  • Read the code your keys flow through — open an issue, send a PR, or fork it.
  • See what we're building — the work in progress lives in our docs and engineering blog.
  • No lock-in — the open-source gateway is MIT-licensed and production-grade. Enterprise adds SSO, RBAC, audit logs, and support on top, not a different core.

Every line is public. The gateway you would run in production is the repo you can read today: issues, PRs, and forks all welcome.

+6
Enterprise support is a dedicated Slack or Teams channel with our engineers, not a ticket queue.

The security controls a review will check for:

Supply chain

Cosign-signed, hardened non-root images — verify provenance before deploy

Scanning

Vulnerability-scanned (Grype, zero high or critical) with CodeQL on the codebase

Identity

SSO, JWT auth, and RBAC through your identity provider, with SCIM provisioning

Guardrails

PII masking and prompt-injection guardrails (Presidio, Lakera, and more)

Secrets

Secrets pulled from AWS Secrets Manager, Vault, or Azure Key Vault, never hardcoded

Residency

Self-hosted with no telemetry, so your data never leaves your infrastructure

Data centre server hall
08

Run it yourself. Start today.

Deploy the open-source gateway in an afternoon, with spend tracked and capped on every request. Add SSO, audit logs, and an SLA when it goes org-wide.