LiteLLM: Traffic Mirroring, Batch Completions, and Traffic to Two Providers
We’re getting ready to launch a self-hosted LLM, and at the testing stage the general idea is to send client requests simultaneously both to the “default production model” like GPT-5.6 and to the model running on our own server. And after getting the responses, we’ll compare them with Phoenix or Opik, and gradually tune our…
https://rtfm.co.ua/en/litellm-traffic-mirroring-batch-completions-and-traffic-to-two-providers/
#AI #LiteLLM #observability #OpenTelemetry #VictoriaTraces
Post #44
25