GUIDE: A Dataset and Study of Context for Intent Recognition in Asymmetric Guiding Dialogue
Open Access
- Author:
- Lyu, Jincheng
- Area of Honors:
- Computer Science
- Degree:
- Bachelor of Science
- Document Type:
- Thesis
- Thesis Supervisors:
- Rebecca Jane Passonneau, Thesis Supervisor
Rebecca Jane Passonneau, Thesis Honors Advisor
Sencun Zhu, Faculty Reader - Keywords:
- Intent recognition
situated dialogue
instructional dialogue
asymmetric dialogue
human–AI interaction
intent classification
dialogue acts
contextual modeling
VirtualHome
multimodal simulation
BERT
pre-trained language models
context window
supervised learning
task-oriented dialogue
dataset creation
annotation quality
input engineering
token representation - Abstract:
- Intent recognition in situated, goal-directed dialogue remains challenging because utterances are highly context dependent and speakers often possess asymmetric knowledge. This thesis stud- ies intent classification in instructional interactions by introducing GUIDE (Growing Understand- ing through Interactive Daily Experience), a curated dataset of 447 transcripts collected from VirtualHome-based simulations of 10 household activities, each instantiated in four controlled variants to elicit guidance, confusion, and error correction. GUIDE includes fine-grained, role- sensitive intent labels (e.g., Inform, Request Information, Request Action, Casual, Greetings), distinguishing human guidance from the alien learner’s knowledge reports and confirmations. We first establish that label quality is pivotal: manual standardization of intents across 23{,}195 utterances improves BERT-base accuracy from 0.64 (uncleaned) to 0.70 (cleaned) under identical settings. We then evaluate pre-trained models (ALBERT, T5, BERT-base, BERT-large) and analyze the effect of context size and input construction. Short multi-turn context (target + 1 prior turn) consistently helps, raising accuracy to 0.74; end padding stabilizes performance (0.75); compact role encoding and input pruning further reduce token overhead; and summing token embeddings across utterances yields a best fine-grained accuracy of 0.76. Aggregating labels into high-level categories substantially simplifies the task, achieving up to 0.87 with BERT-large (1 epoch). Our findings show that (i) careful curation, (ii) short, local context, and (iii) lightweight input engineering materially benefit intent recognition in instructional dialogues. The resulting single- stage, supervised classifier operates offline and does not perform dialogue management or action generation. We conclude with limitations (domain scope, transcript variability) and outline direc- tions for multimodal integration, hierarchical/online classification, and broader data collection.
Accessible Version in Progress
We're generating an accessible version of this file to meet ADA Title II requirements. This process may take up to one hour. Please return later to access the accessible copy once it's ready.
You can still download the current version by clicking "OK".
What's happening:
An accessible PDF is being generated using Adobe with AI used to generate alternative text (alt text) for images in the PDF.