LLM-Advisor: Dynamic Model Selection and Query Routing in Heterogeneous Multi-LLM Architectures

Authors

  • Harshil Lodhiya Chief Software Architect, Product and AI Engineering, SlicedHealth Inc., Vadodara, India

DOI:

https://doi.org/10.65138/ijresm.v9i7.3488

Abstract

The rapid proliferation of Large Language Models (LLMs) with varying capability profiles, context window limits, execution latencies, and financial costs presents a significant operational challenge for enterprise AI deployments. Monolithic deployment strategies wherein all requests are directed to a single high-capability frontier model result in substantial compute over-provisioning and excessive operational costs for routine queries. Conversely, relying solely on lightweight models degrades output accuracy on complex multi-step reasoning tasks. To resolve this trade-off, this paper introduces LLM-Advisor, an open-source, adaptive framework designed for intelligent query categorization, dynamic model evaluation, and constraint-aware request routing across heterogeneous multi-LLM pools. LLM-Advisor analyzes incoming prompt features, structural complexity, domain requirements, and user-defined constraints (e.g., maximum cost per request, latency thresholds) to route tasks to the optimal candidate model. We evaluate LLM-Advisor using a benchmark suite of 1,000 queries across code generation, general reasoning, and contextual retrieval tasks using both proprietary and open-weight models (including GPT-4o, Claude 3.5 Sonnet, Llama 3, and Mistral). Experimental results demonstrate that LLM-Advisor achieves a 42% reduction in overall inference expenditure and a 35% decrease in average response latency while retaining 94.6% task accuracy compared to static GPT-4o baseline routing. These findings highlight LLM-Advisor as an efficient, highly scalable middleware solution for production-grade AI system deployments.

Downloads

Download data is not yet available.

Downloads

Published

23-07-2026

Issue

Section

Articles

How to Cite

[1]
H. Lodhiya, “LLM-Advisor: Dynamic Model Selection and Query Routing in Heterogeneous Multi-LLM Architectures”, IJRESM, vol. 9, no. 7, pp. 84–86, Jul. 2026, doi: 10.65138/ijresm.v9i7.3488.

Most read articles by the same author(s)