Respan routes every LLM call through one gateway, then logs, evaluates and caps spend on it: 500+ models behind one API.
The team behind Keywords AI rebuilt it into an LLM engineering platform. Send OpenAI-style calls through Respan to reach 500+ models, get each request as a trace with per-span latency, set hard spend caps with Slack or email alerts, and fall back to another model when one rate-limits. Prompts and rollouts stay version-controlled from the UI.
Strongest bit: observability and evals sit where the traffic already flows, so you stop gluing LangSmith, a proxy and a billing guard into one stack.
https://respan.ai/
Post #1268
215