Skip to content

Add experimental token velocity aware autoscaling #1546

Description

@asm582

What would you like to be added: Using metrics available in llmd router enable experimental token velocity aware autoscaling for pd and and non-pd deployments

Why is this needed: Inspired by this paper: https://arxiv.org/abs/2512.03416 traditional metrics like queue size etc are lagging indicators that may not be better signals for autoscaling in certain scenario's. Hence, we want to enable autoscaling based on in-flight tokens, a feature exposed by the LLM router that can be used for autoscaling in both PD and non-PD setups, in combination with the KV utilization metric.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions