<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd"><dc:title>GUIDE: A Dataset and Study of Context for Intent Recognition in Asymmetric Guiding  Dialogue</dc:title><dc:creator>Lyu, Jincheng </dc:creator><dc:subject>Intent recognition</dc:subject><dc:subject>situated dialogue</dc:subject><dc:subject>instructional dialogue</dc:subject><dc:subject>asymmetric dialogue</dc:subject><dc:subject>human–AI interaction</dc:subject><dc:subject>intent classification</dc:subject><dc:subject>dialogue acts</dc:subject><dc:subject>contextual modeling</dc:subject><dc:subject>VirtualHome</dc:subject><dc:subject>multimodal simulation</dc:subject><dc:subject>BERT</dc:subject><dc:subject>pre-trained language models</dc:subject><dc:subject>context window</dc:subject><dc:subject>supervised learning</dc:subject><dc:subject>task-oriented dialogue</dc:subject><dc:subject>dataset creation</dc:subject><dc:subject>annotation quality</dc:subject><dc:subject>input engineering</dc:subject><dc:subject>token representation</dc:subject><dc:coverage>Computer Science</dc:coverage><dc:relation>B S</dc:relation><dc:description>Intent recognition in situated, goal-directed dialogue remains challenging because utterances
are highly context dependent and speakers often possess asymmetric knowledge. This thesis stud-
ies intent classification in instructional interactions by introducing GUIDE (Growing Understand-
ing through Interactive Daily Experience), a curated dataset of 447 transcripts collected from
VirtualHome-based simulations of 10 household activities, each instantiated in four controlled
variants to elicit guidance, confusion, and error correction. GUIDE includes fine-grained, role-
sensitive intent labels (e.g., Inform, Request Information, Request Action, Casual, Greetings),
distinguishing human guidance from the alien learner’s knowledge reports and confirmations.
We first establish that label quality is pivotal: manual standardization of intents across 23{,}195
utterances improves BERT-base accuracy from 0.64 (uncleaned) to 0.70 (cleaned) under identical
settings. We then evaluate pre-trained models (ALBERT, T5, BERT-base, BERT-large) and analyze
the effect of context size and input construction. Short multi-turn context (target + 1 prior turn)
consistently helps, raising accuracy to 0.74; end padding stabilizes performance (0.75); compact
role encoding and input pruning further reduce token overhead; and summing token embeddings
across utterances yields a best fine-grained accuracy of 0.76. Aggregating labels into high-level
categories substantially simplifies the task, achieving up to 0.87 with BERT-large (1 epoch).
Our findings show that (i) careful curation, (ii) short, local context, and (iii) lightweight input
engineering materially benefit intent recognition in instructional dialogues. The resulting single-
stage, supervised classifier operates offline and does not perform dialogue management or action
generation. We conclude with limitations (domain scope, transcript variability) and outline direc-
tions for multimodal integration, hierarchical/online classification, and broader data collection.</dc:description><dc:contributor>Rebecca Jane Passonneau, Thesis Supervisor</dc:contributor><dc:contributor>Rebecca Jane Passonneau, Thesis Honors Advisor</dc:contributor><dc:contributor>Sencun Zhu, Faculty Reader</dc:contributor><dc:rights>open_access</dc:rights><dc:date>2025-11-22T15:56:05Z</dc:date><dc:identifier>https://honors.libraries.psu.edu/catalog/9898jjl6992</dc:identifier></oai_dc:dc>