For the complete documentation index, see llms.txt. Markdown versions of all docs pages are available by appending .md to any docs URL.
Inference workloads
Route to your own self-hosted generative AI models with inference workloads.
Inference routing
Route AI inference requests to LLM workloads using the Kubernetes Gateway API Inference Extension.
Multiple inference pools
Route inference requests to multiple InferencePools based on the model name in the request body.
Benchmarking
Explore performance benchmarks, methodology, and reproduction guidance for agentgateway inference …