Skip to content

Cache-aware LLM Routing

When you run a large language model across multiple replicas, it matters a lot which replica a request is sent to. If a replica already has a warm cache for that request's prompt prefix, the Time-To-First-Token and overall response time are significantly faster. Cache-aware LLM routing routes each request to the replica most likely to already have that prompt cached, instead of spreading requests evenly across replicas.

Besides faster responses, this also means your hardware is used more efficiently: fewer requests result in a "cold" generation that recomputes a prefix another replica had already cached.

Cache-aware LLM routing works on top of your service as normal. All regular UbiOps features remain available, such as authentication, rate limiting, request logging and monitoring.

How it works

Cache-aware LLM routing can be enabled with a single toggle while configuring the service. There are no additional configuration parameters to set.

Availability

This feature needs to be enabled for your organization's subscription, and may not be available on every installation. If you don't see the toggle, reach out to your cluster administrator.

Under the hood, when the toggle is enabled, UbiOps deploys llm-d, an open source project for LLM inferencing, to handle routing for that service. Instead of requests being load balanced directly to the deployment replicas, they are first routed through llm-d, which picks the replica expected to give the best cache match for the incoming request, based on previous requests it has observed.

This is fully managed by UbiOps: no separate deployment or configuration is required from you beyond the toggle.

Things to be aware of

Cache-aware routing is most beneficial when:

  • Your service is backed by an LLM deployment running with multiple replicas.
  • It's expected that many requests will share a common prefix. Cache-aware LLM routing can only speed up requests when previous, similar requests have already be routed to one of the LLM replicas.

Cache-aware routing should not be used when:

  • Running a single replica. In this case there is nothing to route between, all requests already go to the same replica.
  • Running non-LLM workloads. This feature is specifically designed around LLM prompt caching behavior and is not useful for services that aren't serving large language models.

It is also good to realize that the llm-d is a highly configurable piece of software and it is not a one-size-fits-all solution for optimal routing. UbiOps deploys llm-d with routing strategies that will improve routing for general cases, but it will not be tuned to your specific setup. For example, the routing is done based on an approximation of the cache contents of the given replicas, it does not use actual cache metrics exposed by the model server (e.g. vLLM). In case a more specific configuration is required, please contact our support team.