<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd"><dc:title>Characterizing Situational Control Assessment Mechanisms for Emotion Prediction Tasks in a Large Language Model</dc:title><dc:creator>Craig, Lucas </dc:creator><dc:subject>large language model</dc:subject><dc:subject>emotion</dc:subject><dc:subject>cognition</dc:subject><dc:subject>explainable ai</dc:subject><dc:subject>mechanistic interpretability</dc:subject><dc:subject>theory of mind</dc:subject><dc:subject>ai</dc:subject><dc:subject>machine learning</dc:subject><dc:coverage>Computer Science</dc:coverage><dc:relation>B S</dc:relation><dc:description>As Large language models (LLMs) have become better at imitating and understanding human behaviors, their ability to predict human emotional responses to situations has improved. Prior work has shown that these predictions are influenced by cognitive appraisal dimensions, of the kind proposed by the appraisal theory of emotion in psychology. More recent studies have identified the general location of the appraisal within LLMs, but our understanding of how LLMs perform appraisals remains fairly limited. In this thesis, I take a step towards bridging this gap by studying a small but capable LLM: Qwen3-4B. Using a set of minimally-differing prompts, along with various tools from the “circuit analysis” sub-field of AI interpretability, I attempt to map out the mechanisms responsible for using human-controllable vs. situational appraisals of negative events to inform emotion predictions. I identify a circuit that modulates the anger-vs-sadness distinction on a negative-valence emotion prediction task; testing the circuit across a broad set of agent types reveals that it does not cleanly implement human controllability as defined by appraisal theory. Instead, it appears to detect something like autonomous animate agency, modulated by intentionality modifiers (e.g. "deliberately," "accidentally"). These results suggest that Qwen3-4B has learned a situational control-adjacent appraisal dimension, one which also bundles together distinctions that psychological appraisal theory keeps separate  (e.g. human vs. animal agency, intentionality).</dc:description><dc:contributor>James Z Wang, Thesis Supervisor</dc:contributor><dc:contributor>Sencun Zhu, Thesis Honors Advisor</dc:contributor><dc:contributor>Shagufta Mehnaz, Faculty Reader</dc:contributor><dc:rights>open_access</dc:rights><dc:date>2026-04-19T20:58:29Z</dc:date><dc:identifier>https://honors.libraries.psu.edu/catalog/10329lkc5576</dc:identifier></oai_dc:dc>