Control & Visibility

Cloudflare's AI Gateway allows you to gain visibility and control over your AI apps. By connecting your apps to AI Gateway, you can gather insights on how people are using your application with analytics and logging and then control how your application scales with features such as caching, rate limiting, as well as request retries, model fallback, and more.

Strategic Value Pillars

Observability

Monitor your AI ecosystem in real-time. Gain insights into request volume, token consumption patterns, and precise inference costs across all connected providers.

Scalability & Optimization

Enforce operational efficiency. Implement semantic caching to slash latency and token costs, and apply rate limiting to protect upstream provider quotas.

Application Resiliency

Eliminate provider downtime. Automatically retry failed inference calls or trigger model fallbacks to ensure constant availability for your end-users.

Topological Comparison Matrix

Toggle between configurations to evaluate the difference between direct provider connection and routing via Cloudflare AI Gateway. Click the β“˜ badges on each node for a plain-English + technical explanation.

Advanced Gateway Configuration

Configure headers and dashboard policies to enforce technical governance over your inference workloads.

Bring Your Own Keys (BYOK)

Securely vault your AI provider keys within Cloudflare Secrets Store. Eliminate hardcoded keys in application code and pass references via the gateway configuration.

Custom Costs Management

Inject accurate spend metrics. Use the cf-aig-custom-cost header to track real ROI in your analytics dashboard based on negotiated enterprise rates.

Request Handling & Resiliency

Define cf-aig-request-timeout and retry headers (cf-aig-max-attempts) to handle provider lag or temporary service degradation seamlessly.

Gateway Authentication

Harden your proxy. Require a cf-aig-authorization header (Cloudflare API Token) to ensure only authorized clients can route through your gateway.

Custom Providers

Extend Gateway features (Caching, Logs) to regional endpoints or internal HTTPS-based models by defining a custom provider base URL and slug.

Management & Logs

Programmatically create gateways via API. Manage log retention and export audit trails to external SIEMs for enterprise compliance (Logpush).

Models & Autonomous Agents

A standardized control plane to orchestrate authentication, metrics, and routing across a heterogeneous matrix of foundation models and AI Agents.

Universal Provider Support

AI Gateway supports Universal Routing to major providers including Cloudflare Workers AI, OpenAI, Anthropic, Google Vertex AI, and more through a unified schema.

Agent Setup & Tracing

Autonomous AI Agents rely on multi-step reasoning. AI Gateway logs every tool call and reasoning step, providing full observability into complex agentic loops and autonomous actions.

Standardized Client Integration

Point your existing SDK's base URL to the AI Gateway service layer. Zero refactor required for core logic.

Python Integration (OpenAI SDK)

from openai import OpenAI client = OpenAI( api_key="provider-api-key", # Route via Cloudflare AI Gateway base_url="https://gateway.ai.cloudflare.com/v1/{account_id}/{gateway_id}/openai" )

Workers AI Native Implementation

const response = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", { prompt: "Initiate Security Audit" }, // Define the gateway binding options { gateway: { id: "my-gateway-id" } } );

Operational Visibility Lab

Simulate requests to observe the logging output, caching states, and failover mechanics exactly as recorded by the gateway.

nanosek@cf-ai-gateway-lab:~
>_
Awaiting stateful transaction telemetry...

Live AI Threat Defense

Select an attack vector and see how Cloudflare AI Gateway intercepts threats at the edge before they hit your LLM and drain your wallet.

πŸ‘Ύ Threat Actor
πŸ›‘οΈ AI Gateway
🧠 Upstream LLM
Edge Telemetry Stream
[SYSTEM] Ready. Select attack type and Defense mode.

Node

πŸ‘‹ In plain English

βš™οΈ Technical detail