nexus-llm-router
Intelligent multi-LLM routing middleware with task-aware model selection, cost optimization, fallback safety, and a drop-in OpenAI-compatible API.

Why Nexus
Most teams start with one LLM endpoint. That works until traffic grows, latency starts swinging, finance asks why every request hits the most expensive model, and incident review asks why the app kept calling a degraded provider. Nexus gives the application one stable OpenAI-compatible API while moving model choice, fallback, budget, audit, and routing rationale into infra-owned middleware.
Nexus is designed for AI infrastructure engineers running multi-model production pipelines where quality, latency, and cost must be optimized at the same time.
Problems It Solves
-
Issue: every prompt is sent to the same frontier model. Nexus solves this by classifying prompt complexity and routing simple tasks to cheaper low-latency models while reserving premium models for hard prompts.
-
Issue: spend grows faster than product usage. Nexus solves this with cost-aware routing, model cost estimates, per-user budget guardrails, and Prometheus cost metrics.
-
Issue: code, medical, legal, and general prompts need different quality defaults. Nexus solves this by extracting a domain tag and applying deterministic policy rules such as medical/legal to Claude Sonnet 4.6 and complex code to GPT-5.5.
-
Issue: one provider has an incident and the app fails hard. Nexus solves this with per-provider circuit breakers and automatic fallback chains.
-
Issue: provider latency spikes during peak traffic. Nexus solves this with latency-aware routing that tracks rolling p95 latency and penalizes slow providers.
-
Issue: teams want to compare models without rewriting product code. Nexus solves this with stable request-id A/B routing selected by the
X-Router-Strategyheader. -
Issue: support and compliance teams ask why a model answered a request. Nexus solves this by persisting durable audit records with
request_id, selected model, strategy, rationale, latency, token usage, and cost. -
Issue: a single API key can overwhelm the router. Nexus solves this with a token-bucket rate limiter keyed by API key identifier.
-
Issue: session or tenant budgets need hard enforcement. Nexus solves this by rejecting requests before dispatch when estimated spend would exceed the configured cap.
-
Issue: PII can leak into third-party providers. Nexus solves this with optional regex redaction and a Presidio extension path before provider dispatch.
-
Issue: teams need OpenAI compatibility without giving up provider choice. Nexus solves this by exposing
/v1/chat/completionswhile normalizing OpenAI, Anthropic, Gemini, and Moonshot adapters behind one interface. -
Issue: model routing becomes a hidden product decision. Nexus solves this by making routing policy explicit, testable, observable, and owned in infra.
Demo Gallery
Terminal routing demo with JSON rationale logs:

Observe -> Decide -> Act state-machine demo:

Prompt-prefix cache affinity demo:

Soft rate-limit avoidance demo:

Features
- Router engine with configurable strategies
- Adapter pipeline with full observability
- Async-first design using
asyncio+httpx - Type-safe with full
mypycompliance - Production-ready with Docker, CI/CD, and structured logging
Quick Start
git clone https://github.com/Francis1998/nexus-llm-router.git
cd nexus-llm-router
pip install -e ".[dev]"
cp .env.example .env
PYTHONPATH=src uvicorn api.main:app --reload
Quality Gates
ruff check src/ tests/ scripts/
mypy src/
pytest tests/ -v
Docker Compose
docker compose up --build
Services:
- Router:
http://localhost:8000 - Prometheus:
http://localhost:9090 - Grafana:
http://localhost:3000
Routing Strategies
Select a strategy with X-Router-Strategy:
rule-based: domain and complexity priority matrixclassifier: logistic-regression-style complexity and domain featurescost-optimal: minimizes estimated cost subject to quality floorlatency-aware: penalizes providers with poor rolling p95 latencyreliability-aware: routes to the highest-quality model whose provider circuit is closed, and orders the fallback chain healthy-providers-firstweighted-blend: selects the model with the highest tunable composite of normalized quality, cost, and latency (weights viaNEXUS_BLEND_*)budget-aware: selects the highest-quality model whose estimated per-request cost stays within a hard ceiling (NEXUS_REQUEST_COST_CEILING_USD); the dual ofcost-optimalprovider-family-cost-ceiling: selects the highest-quality domain-eligible model whose estimated cost stays within the ceiling for its provider family (openai/anthropic/google/moonshot); default viaNEXUS_PROVIDER_FAMILY_COST_CEILING_USD, with cross-family fallback when nothing fits — OpenRouter/LiteLLM-style family budgets for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2sticky-session: consistent-hashessession_idonto one domain-eligible model, so every turn in a session routes to the same model (context/prompt-cache affinity) while distinct sessions spread across the poolsticky-tenant-hash: consistent-hashesmetadata.tenant_id(thenuser_id/sticky_keyfallbacks) onto one domain-eligible model per tenant with healthy ring failover — distinct fromsticky-session, which pins onlysession_idfor conversational affinity across GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 trafficvalue: selects the model with the best quality-per-dollar ratio, maximizing spend efficiency with no threshold to tunecanary: rolls a configurable traffic fraction (NEXUS_CANARY_WEIGHT) onto a canary model (NEXUS_CANARY_MODEL) while the rest stays on a stable model (NEXUS_CANARY_STABLE_MODEL); health-gated, so a canary whose provider circuit is open is paused and all traffic falls back to the stable modelcanary-tier-blend: blends canary traffic with complexity-tier affinity — on the canary slice prefer the canary when it matches the inferred tier, else canary; off-slice or when unhealthy prefer tier match, else highest quality (NEXUS_CANARY_*)shadow-traffic-mirror: cost-optimal primary routing (NEXUS_QUALITY_FLOOR) with a deterministicrequest_idslice (NEXUS_SHADOW_TRAFFIC_PERCENT, default5) that annotates a shadow mirror model from a different provider for dual-run telemetry — LiteLLM/OpenRouter-style shadow comparison without changing the returned primarycanary-cost-blend: blends cost exploration with healthy-provider minimization — default picks the cheapest healthy model, whileNEXUS_CANARY_COST_BLEND_PERCENT(default10) explores the next-cheaper healthy tier via deterministicrequest_idhashing; distinct fromcanary-tier-blendtoken-cost-anomaly-shed: sheds to cheaper healthy models when the top quality pick's projected cost/1k exceeds the rolling mean timesNEXUS_TOKEN_COST_ANOMALY_RATIO(default2.0); falls back to quality ranking when no cheaper healthy option exists — LiteLLM/OpenRouter-style spend spike guardrails for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2multi-region-latency-hedge: stays on highest-quality primary-region models but hedges to the lowest-p50 secondary-region candidate when the primary provider p50 exceedsNEXUS_LATENCY_HEDGE_MS(default500) — regional latency escape hatch for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2adaptive-timeout-hedge: keeps the highest-quality eligible model unless its rolling provider p95 exceeds the fastest observed eligible p95 byNEXUS_ADAPTIVE_TIMEOUT_HEDGE_RATIO(default1.5), then hedges to the fastest observed alternative for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2token-bucket-tenant: maintains independent tenant request-token buckets (NEXUS_TOKEN_BUCKET_TENANT_RATE, default5/s); in-budget requests keep quality-first routing while over-budget requests shed to the cheapest domain-eligible GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 modelregion-carbon-blend: blends regional carbon intensity with rolling provider p95 latency viaNEXUS_REGION_CARBON_BLEND_WEIGHT(default0.5;0= latency only,1= carbon only) for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2provider-weight-decay: exponentially decays provider selection weight after failures (NEXUS_PROVIDER_WEIGHT_DECAY_FACTOR, default0.5) and recovers slowly on success (NEXUS_PROVIDER_WEIGHT_RECOVER, default0.1) for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2retry-after-respect: skips providers still inside a Retry-After cooldown (NEXUS_RETRY_AFTER_DEFAULT_SECONDS, default30) and falls back to the next healthy provider for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2latency-slope-shed: sheds the quality leader when its EWMA latency slope exceedsNEXUS_LATENCY_SLOPE_THRESHOLD_MS(default25ms/step; window viaNEXUS_LATENCY_SLOPE_WINDOW, default10) to a lower-latency / cheaper healthy model for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2latency-budget: selects the highest-quality model whose provider rolling p95 latency stays within a hard SLA (NEXUS_LATENCY_SLA_MS); the latency-domain dual ofbudget-aware, trading quality for speed only when the SLA requires itprompt-length-tier-shed: sheds frontier-tier models whenprompt_tokens_estimateexceedsNEXUS_PROMPT_LENGTH_TIER_TOKENS(default8000) and picks the best mid/economy alternative; short prompts keep pure quality ranking — LiteLLM/OpenRouter-style length tier shedding for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2retry-budget-aware-failover: prefers highest-quality healthy models whilemetadata.retry_remaining(orNEXUS_RETRY_BUDGET_DEFAULT, default3) is > 1, then failovers to lowest-latency healthy model on the last attempt — LiteLLM/OpenRouter-style retry-budget routing for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2cache-hit-sticky-warm-pool: consistent-hashes a long prompt prefix (minNEXUS_CACHE_HIT_STICKY_MIN_CHARS, default64) onto one domain-eligible model with healthy ring failover so provider prompt caches stay warm across GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 trafficembedding-cache-key-namespace: consistent-hashes{NEXUS_EMBEDDING_CACHE_NAMESPACE_PREFIX}:{tenant}(default prefixembed) onto one domain-eligible model with healthy ring failover so embedding/cache keys stay isolated across tenants for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 trafficcarbon-aware-preference: prefers lower carbon-intensity providers underNEXUS_CARBON_AWARE_MAX_INTENSITY(default400) for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2tenant-concurrency-lease: prefers providers with remaining per-tenant in-flight headroom (NEXUS_TENANT_CONCURRENCY_LEASE, default8) usingInflightStatskeyed by tenant/session — LiteLLM/Portkey-style tenant fairness for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2provider-error-budget-shed: prefers healthy domain-eligible providers whose rollingSuccessStatserror rate stays underNEXUS_PROVIDER_ERROR_BUDGET_RATE(default0.15); when every provider is over budget it falls back to lowest error rate, then quality, for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 trafficregion-latency-p99-shed: prefers region-matching domain-eligible providers whose rollingLatencyStatsp99 stays underNEXUS_REGION_LATENCY_P99_MS(default3000); when every regional provider is hot it falls back to lowest p99, then quality — LiteLLM/OpenRouter-style regional tail shedding for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2sticky-canary-cost: pins tenants via consistent hashing onmetadata.tenant_id(with user/session fallbacks) and blends a deterministicrequest_idexplore slice (NEXUS_STICKY_CANARY_COST_PERCENT, default10) toward cheaper healthy models while keeping sticky affinity off-slice — LiteLLM/Portkey-style sticky cost canaries for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2queue-depth-fairness: prefers domain-eligible providers whose liveInflightStatsqueue depth stays underNEXUS_QUEUE_DEPTH_SOFT_CAP(default4); when every provider is saturated it falls back to lowest depth, then quality — LiteLLM/Portkey-style queue fairness for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2provider-quota-fair-share: tracks the lastNEXUS_PROVIDER_QUOTA_LOOKBACKselections (default100) and prefers eligible providers below equal request share, shedding over-share providers while preserving quality/cost tie-breaks for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2provider-spend-telemetry: prefers lower-spend providers once soft USD spend telemetry exceedsNEXUS_PROVIDER_SPEND_SOFT_USD(default10) for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 fleetssemantic-cache-ttl-affinity: pins cacheable requests to providers with warm TTL underNEXUS_SEMANTIC_CACHE_TTL_SECONDS(default300) for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2circuit-breaker-half-open-probe: prefers healthy closed providers and allows only limited concurrent probes into half-open/recovering providers (NEXUS_CIRCUIT_HALF_OPEN_PROBE_BUDGET, default2) — LiteLLM/Portkey-style half-open probe budgeting for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2latency-slo-shed: sheds providers whose rolling p95 exceedsNEXUS_LATENCY_SLO_MS(default2000) when faster alternatives exist; prefers highest quality among under-SLO candidates and falls back to lowest latency when every provider is hot — LiteLLM/OpenRouter-style latency SLO shedding for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2adaptive-timeout: selects the highest-quality model whose risk-adjusted provider p95 fits an adaptive timeout budget derived from request urgency, recent latency, and success/error signals; prefers faster models under realtime pressure and allows slower higher-quality models when comfortablecomplexity-tier: treats the classifier complexity score as a required quality target and picks the cheapest model meeting it — a catalog-adaptive quality-for-cost escalation ladder with no thresholds to tune (falls back to the top-quality model when the target is unreachable)round-robin: load-balances across every provider offering a domain-eligible model (routing each to that provider's best eligible model), spreading rate-limit pressure instead of converging on one provider; balanced by a stablerequest_idhash so routing stays deterministic and replayablecascade: routes the primary attempt to the cheapest domain-eligible model and orders the fallback chain by ascending cost, so a failure escalates one price/capability rung at a time instead of jumping to the top-quality model — minimizing expected spend on the common first-attempt-succeeds path with no thresholds to tuneepsilon-greedy: with probabilityNEXUS_EPSILONexplores by picking uniformly among domain-eligible models (stable second hash ofrequest_id); otherwise exploits the highest-quality eligible model — a replayable bandit policy so under-prioritized catalog entries still get live trafficadaptive-exploration: likeepsilon-greedy, but epsilon decays fromNEXUS_ADAPTIVE_EXPLORATION_BASE(default0.2) towardNEXUS_ADAPTIVE_EXPLORATION_MIN(default0.02) asSuccessStatstotal successes grow — explore more while cold, exploit more as GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 traffic proves outgeo-region: prefers models whosesupported_regionsinclude the request region (GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 catalog priors)region-tier-affinity: prefers models matching both request geo region and complexity-mapped tier (frontier/mid/economy), then tier, then region, then quality — no extraNEXUS_*knobssoft-family-budget: deprioritizes provider families whose rolling observed spend exceeds a soft budget (NEXUS_SOFT_FAMILY_BUDGET_USD, window viaNEXUS_SOFT_FAMILY_BUDGET_WINDOW_SECONDS); prefers highest-quality models from under-budget families and falls back to the cheapest other family when every family is hot — OpenRouter/LiteLLM-style family spend steering for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2sticky-region-failover: pinssession_idto a model inside the first healthy preferred region (requestregionfirst, thenNEXUS_STICKY_REGION_FAILOVER_PREFERENCES), failovers to the next region when the preferred pool is unhealthy, and keeps sticky affinity when healthy — geo-residency plus session stickiness for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2region-failover-hysteresis: like sticky-region-failover but waits forNEXUS_REGION_FAILOVER_HYSTERESIS_SUCCESSES(default3) consecutive preferred-region successes before flapping back after a failover for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2token-budget: selects the highest-quality domain-eligible model whosecontext_windowfitsprompt_tokens_estimate + max_tokenswithin the requesttoken_budget; falls back to the largest-context model when nothing fitsslo-aware: selects the highest-quality domain-eligible model whose provider rolling success rate meetsNEXUS_AVAILABILITY_SLO; falls back to the highest success-rate model when nothing meets the SLOsemantic-cache: onmetadata.cache_hit, prefers the cheapest domain-eligible model; on miss, falls through to cost-optimal under the quality floorleast-busy: selects the highest-quality domain-eligible model on the provider with the lowest current in-flight load; load ties prefer higher quality, then lower estimated costprompt-prefix-cache: hashes long shared system-prompt prefixes to sticky provider/model buckets, improving OpenRouter/LiteLLM-style KV-cache affinity for GPT-5.5, Claude Sonnet 4.6, Gemini 2.5, and Kimi K2; short prefixes fall back to cost-optimalconcurrency-cap: skips providers whose live in-flight count is at or aboveNEXUS_CONCURRENCY_CAP, then selects the highest-quality remaining model for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 trafficprompt-prefix-cache: hashes long shared system-prompt prefixes to sticky provider/model buckets, improving OpenRouter/LiteLLM-style KV-cache affinity for GPT-5.5, Claude Sonnet 4.6, Gemini 3.x, and Kimi K2; short prefixes fall back to cost-optimalsoft-rate-limit: prefers healthy providers with fewer recent 429/rate-limit observations, so GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 traffic backs away from quota pressure before hard circuit breakers tripcost-latency-pareto: keeps non-dominated cost/latency candidates (Pareto front on estimated spend and rolling provider p95), then breaks ties by quality — LiteLLM/Portkey-style multi-objective routing across GPT-5.5, Claude Sonnet 4.6, Gemini 3.x, and Kimi K2token-bucket-burst: maintains per-provider token buckets (NEXUS_TOKEN_BUCKET_CAPACITY,NEXUS_TOKEN_BUCKET_REFILL_PER_SEC) and prefers providers with burst quota for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 traffic; when every bucket is empty it falls back to the highest remaining fraction, then costmodel-tier-rate-limit: infers frontier/mid/economy tiers from model names and applies tier-specific soft RPM ceilings per provider (NEXUS_TIER_FRONTIER_RPM,NEXUS_TIER_MID_RPM,NEXUS_TIER_ECONOMY_RPM); prefers providers under their tier limit for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 traffic and falls back to the least-saturated providerfailover-priority: walks an explicit ordered model preference list and picks the first healthy provider (LiteLLM-style ordered failover)provider-health-score-blend: blends circuit availability, rolling success rate, inverse p95 latency, model quality, and inverse estimated cost; open circuits are skipped whenever a healthy provider exists (NEXUS_HEALTH_BLEND_*)health-cost-latency: ternary blend of rolling provider success rate, inverse estimated cost, and inverse p95 latency for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2 traffic (NEXUS_HCL_*)provider-hourly-cost-ceiling: skips providers whose rolling hourly estimated spend exceedsNEXUS_PROVIDER_HOURLY_COST_CEILING_USD(default5.0), preferring highest quality under ceiling — distinct fromprovider-family-cost-ceilingfor GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2quality-weighted-sticky: sticky-session hashing with hash-ring bucket weights proportional toquality_score(higher quality gets larger sticky share) — distinct from uniformsticky-sessionandsticky-tenant-hashfor GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2token-rpm-ceiling,provider-circuit-probe,carbon-latency-blend: tracks estimated prompt tokens per provider over a rolling 60-second window and sheds requests that would exceedNEXUS_TOKEN_RPM_CEILING(default100000) to the next eligible provider for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2adaptive-concurrency-cap: scales per-provider in-flight caps by rolling success rate and inverse p95 latency (NEXUS_ADAPTIVE_CONCURRENCY_BASE_CAP,NEXUS_ADAPTIVE_CONCURRENCY_MIN_CAP,NEXUS_ADAPTIVE_CONCURRENCY_LATENCY_MS) so unhealthy backends shed load while quality-first routing continues for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2provider-token-fair-share: fair-share prompt-token budget per provider in a rolling 60-second window (NEXUS_PROVIDER_TOKEN_FAIR_SHARE_CEILING, default100000) with round-robin weighted by remaining quota for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2tenant-budget-cascade: tracks per-tenant rolling spend, keeps quality-first choices while projected spend fitsNEXUS_TENANT_BUDGET_CASCADE_SOFT, sheds to cheaper providers up toNEXUS_TENANT_BUDGET_CASCADE_HARD, then fails closed with a clear rationale for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2provider-error-budget-reset: temporarily sheds providers aboveNEXUS_PROVIDER_ERROR_BUDGET_RESET_FRACTIONand automatically restores them afterNEXUS_PROVIDER_ERROR_BUDGET_RESET_SECONDS, distinct from cumulativeprovider-error-budget-shed, for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2sticky-region-warmup: sends each new session's firstNEXUS_STICKY_REGION_WARMUP_REQUESTSrequests to a warmup region, then pins the session to its requested or hash-selected region to prevent cold-start flaps for GPT-5.5 / Claude Sonnet 4.6 / Gemini 3.x / Kimi K2ab: deterministic request-id buckets across two model arms
Documentation
| Document | Description |
|---|---|
| Architecture | System design and component overview |
| Configuration | All configuration options |
| Epsilon-greedy guide | Explore/exploit routing walkthrough |
| Adaptive-exploration guide | Decaying epsilon explore/exploit walkthrough |
| Token-budget guide | Context-window-aware quality routing |
| Geo-region guide | Region/residency-aware model selection |
| Region-tier-affinity guide | Combined geo-region and complexity-tier affinity routing |
| Soft-family-budget guide | Rolling soft spend budgets per provider family |
| Sticky-region-failover guide | Session stickiness with ordered region failover |
| Region-failover-hysteresis guide | Region failover with recovery hysteresis |
| Sticky-tenant-hash guide | Per-tenant consistent hashing with healthy failover |
| Embedding-cache-key-namespace guide | Tenant-isolated embedding/cache sticky namespace routing |
| Semantic-cache-ttl-affinity guide | Warm semantic-cache TTL sticky routing |
| Tenant-concurrency-lease guide | Per-tenant in-flight concurrency lease routing |
| Provider-error-budget-shed guide | Rolling provider error-budget shedding |
| Region-latency-p99-shed guide | Regional p99 tail-latency shedding |
| Sticky-canary-cost guide | Sticky tenant affinity with cost canary blend |
| Queue-depth-fairness guide | Soft queue-depth fairness across providers |
| Provider-quota-fair-share guide | Rolling equal-share provider quota routing |
| Provider-token-fair-share guide | Rolling token fair-share routing weighted by remaining quota |
| Tenant-budget-cascade guide | Per-tenant rolling spend cascade with a hard fail-closed ceiling |
| Circuit-breaker-half-open-probe guide | Half-open recovery probe budget routing |
| SLO-aware guide | Availability-SLO quality routing |
| Adaptive-timeout guide | Timeout-adaptive quality routing |
| Adaptive-timeout-hedge guide | Relative p95 hedge from a quality-first provider choice |
| Token-bucket-tenant guide | Per-tenant request budget with cheapest-model shedding |
| Region-carbon-blend guide | Carbon intensity blended with latency scoring |
| Provider-weight-decay guide | Exponential provider weight decay with slow recovery |
| Retry-after-respect guide | Honor provider Retry-After cooldowns |
| Semantic-cache guide | Cache-hit cheapest / miss cost-optimal routing |
| Least-busy guide | Live in-flight load-aware routing |
| Prompt-prefix-cache guide | Sticky system-prompt prefix affinity for provider KV-cache hits |
| Concurrency-cap guide | Per-provider in-flight saturation cap routing |
| Adaptive-concurrency-cap guide | Health-derived dynamic in-flight cap routing |
| Soft-rate-limit guide | Soft 429/rate-limit pressure avoidance |
| Cost/latency Pareto guide | Multi-objective non-dominated cost + latency routing |
| Token-bucket-burst guide | Bursty per-provider token-bucket quota routing |
| Model-tier-rate-limit guide | Tier-specific soft RPM routing by model name |
| Failover-priority guide | Ordered healthy-provider failover |
| Provider-health score blend guide | LiteLLM/Portkey-style health-aware blended routing |
| Health/cost/latency guide | Ternary health, cost, and latency blend routing |
| Provider-family cost-ceiling guide | Per-provider-family spend ceilings for multi-provider budgets |
| Canary-tier-blend guide | Canary rollout with complexity-tier affinity |
| Latency-SLO-shed guide | Latency SLO shedding with under-SLO quality preference |
| Shadow-traffic-mirror guide | Cost-optimal primary with shadow mirror telemetry |
| Canary-cost-blend guide | Cost exploration with next-cheaper healthy tier sampling |
| Token-cost-anomaly-shed guide | Rolling cost/1k anomaly shedding with quality fallback |
| Quickstart | Local setup and first request |
| Safety | Guardrails, fallback, and PII controls |
| Contributing | Development workflow and PR process |
| Security | Vulnerability reporting policy |
| Changelog | Version history |
License
Apache-2.0 © Francis1998
Last updated: 2026-08-01