Radar van Elk Solutions

AI Engineer · AI

Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI

The routing weights inside OpenAI's inference load balancer used to come out of a feedback loop. Engines reported signals, a controller smoothed them into a score, compared it to the fleet average, and nudged each weight up or down. A proportional controller, Lu Zhang notes, with real virtues: many signals folded into one decision, and constrained engines balanced themselves. It also produced behavior nobody could explain well. Ask why one engine

Introductie van de bron.

AI Engineer