Bandits with Noisy Surrogates: Applications in LLM Routing and Variance Reduction
Restricted (Penn State Only)
- Author:
- Singh, Ronak
- Area of Honors:
- Computer Science
- Degree:
- Bachelor of Science
- Document Type:
- Thesis
- Thesis Supervisors:
- Hadi Hosseini, Thesis Supervisor
Martin Fürer, Thesis Honors Advisor - Keywords:
- Bandit Learning
Multi-Armed Bandits
LLM Routing
Online Learning
Contextual Bandits
Linear Bandits
Variance-Aware Bandits
Model Routing
Surrogate Rewards - Abstract:
- Sequential decision-making problems, such as the routing of queries to Large Language Models (LLMs), often suffer from bandit feedback where only the chosen action's reward is observed. While machine learning-generated surrogate rewards can provide valuable side information about unobserved actions, integrating these predictions naively can lead to suboptimal policies due to inherent bias and noise. This thesis explores how to safely and effectively leverage potentially misspecified ML-generated surrogate rewards to accelerate online learning without sacrificing theoretical guarantees. First, motivated by correlation-aware LLM routing, we introduce CABS-C, a contextual bandit algorithm that utilizes bias estimation and importance-weighting to pool true and surrogate rewards into a single learner model. We demonstrate that when surrogate noise is bounded, CABS-C achieves significantly improved regret guarantees compared to standard baselines. Second, we investigate the variance-aware linear bandit setting, introducing the MLA-VOFUL algorithm. Rather than treating surrogates directly as rewards, this approach utilizes Prediction-Powered Inference (PPI) techniques to employ them as control variates. This structural variance reduction isolates the surrogate's bias, achieving strictly improved asymptotic regret whenever surrogate and true rewards are correlated and the offline training dataset is sufficiently large. Together, these approaches provide a robust framework for harnessing imperfect auxiliary predictions in both discrete, context-driven settings and continuous, variance-adaptive environments.
Accessible Version in Progress
We're generating an accessible version of this file to meet ADA Title II requirements. This process may take up to one hour. Please return later to access the accessible copy once it's ready.
You can still download the current version by clicking "OK".
What's happening:
An accessible PDF is being generated using Adobe with AI used to generate alternative text (alt text) for images in the PDF.