<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd"><dc:title>Bandits with Noisy Surrogates: Applications in LLM Routing and Variance Reduction</dc:title><dc:creator>Singh, Ronak </dc:creator><dc:subject>Bandit Learning</dc:subject><dc:subject>Multi-Armed Bandits</dc:subject><dc:subject>LLM Routing</dc:subject><dc:subject>Online Learning</dc:subject><dc:subject>Contextual Bandits</dc:subject><dc:subject>Linear Bandits</dc:subject><dc:subject>Variance-Aware Bandits</dc:subject><dc:subject>Model Routing</dc:subject><dc:subject>Surrogate Rewards</dc:subject><dc:coverage>Computer Science</dc:coverage><dc:relation>B S</dc:relation><dc:description>Sequential decision-making problems, such as the routing of queries to Large Language Models (LLMs), often suffer from bandit feedback where only the chosen action's reward is observed. While machine learning-generated surrogate rewards can provide valuable side information about unobserved actions, integrating these predictions naively can lead to suboptimal policies due to inherent bias and noise. This thesis explores how to safely and effectively leverage potentially misspecified ML-generated surrogate rewards to accelerate online learning without sacrificing theoretical guarantees. First, motivated by correlation-aware LLM routing, we introduce CABS-C, a contextual bandit algorithm that utilizes bias estimation and importance-weighting to pool true and surrogate rewards into a single learner model. We demonstrate that when surrogate noise is bounded, CABS-C achieves significantly improved regret guarantees compared to standard baselines. Second, we investigate the variance-aware linear bandit setting, introducing the MLA-VOFUL algorithm. Rather than treating surrogates directly as rewards, this approach utilizes Prediction-Powered Inference (PPI) techniques to employ them as control variates. This structural variance reduction isolates the surrogate's bias, achieving strictly improved asymptotic regret whenever surrogate and true rewards are correlated and the offline training dataset is sufficiently large. Together, these approaches provide a robust framework for harnessing imperfect auxiliary predictions in both discrete, context-driven settings and continuous, variance-adaptive environments.</dc:description><dc:contributor>Hadi Hosseini, Thesis Supervisor</dc:contributor><dc:contributor>Martin Fürer, Thesis Honors Advisor</dc:contributor><dc:rights>restricted_to_institution</dc:rights><dc:date>2026-03-24T04:19:37Z</dc:date><dc:identifier>https://honors.libraries.psu.edu/catalog/10263rjs7006</dc:identifier></oai_dc:dc>