What would you like to be added: Using metrics available in llmd router enable experimental token velocity aware autoscaling for pd and and non-pd deployments
Why is this needed: Inspired by this paper: https://arxiv.org/abs/2512.03416 traditional metrics like queue size etc are lagging indicators that may not be better signals for autoscaling in certain scenario's. Hence, we want to enable autoscaling based on in-flight tokens, a feature exposed by the LLM router that can be used for autoscaling in both PD and non-PD setups, in combination with the KV utilization metric.
What would you like to be added: Using metrics available in llmd router enable experimental token velocity aware autoscaling for pd and and non-pd deployments
Why is this needed: Inspired by this paper: https://arxiv.org/abs/2512.03416 traditional metrics like queue size etc are lagging indicators that may not be better signals for autoscaling in certain scenario's. Hence, we want to enable autoscaling based on in-flight tokens, a feature exposed by the LLM router that can be used for autoscaling in both PD and non-PD setups, in combination with the KV utilization metric.