The AI routing tollbooth: one API in, ~5% of inference spend skimmed at a gate between app developers and model labs, Stripe layering payments and usage metering on the same traffic Developers call one API routing sits dev + lab apps agents workflows one API Route to the cheapest/best model ~5% of inference spend inference spend ($) ~5% take Sacra estimate Model labs compete on price and quality for the routed traffic OpenAI Anthropic Google · xAI DeepSeek · more compute too pricing compression feeds the toll Stripe layers payments + metering (Metronome) · charge on token and dollar A fee on every machine-speed transaction owns no model · compute · lab needs the rails above