Assessing the Reliability of Large Language Models in Evaluating Mathematical Interestingness
Open Access
- Author:
- Arora, Varnikaa
- Area of Honors:
- Computer Science
- Degree:
- Bachelor of Science
- Document Type:
- Thesis
- Thesis Supervisors:
- Abhinav Verma, Thesis Supervisor
Rebecca Jane Passonneau, Thesis Honors Advisor - Keywords:
- Large Language Models
Mathematical Interestingness
Automated Mathematical Discovery
Jargon Illusion
Semantic Equivalence
Linguistic Bias
Prompt Sensitivity
Heuristics
Neurosymbolic Architecture
GPT-4o
DeepSeek-V3
Large Language Models
Mathematical Interestingness
Automated Mathematical Discovery
Jargon Illusion
Semantic Equivalence
Linguistic Bias
Prompt Sensitivity
Heuristics
Neurosymbolic Architecture
GPT-4o
DeepSeek-V3 - Abstract:
- The question of what makes a mathematical problem interesting is important to both mathe- matics education and the field of automated mathematical discovery (AMD). As large language models (LLMs) are increasingly used to evaluate, generate, and prioritize mathematical content, understanding whether different models apply consistent and coherent criteria for interesting- ness becomes an important factor. This thesis investigates whether LLMs (OpenAI GPT-4o and DeepSeek-V3) give comparable ratings of interestingness when evaluating identical mathematical problems, and tests the structural integrity of their underlying evaluative rubrics. Using a structured prompting methodology with a reproducible random seed, both models were applied to 100 problems drawn from the 42 math dataset. Initially, the models agreed exactly on only 44% of problems, showing substantial divergence on 24% of problems (up to a 5-point difference on a 10-point scale). Analysis revealed this divergence came from conflicting default heuristics. DeepSeek applied a strict novelty-and-non-obviousness standard, which resulted in a strongly bimodal distribution, while OpenAI correlated foundational, historical significance with interestingness. However, to determine if these differing heuristics were grounded in actual mathematical com- prehension, a ”Semantic Equivalence” test was conducted. Ten mathematically trivial problems were rewritten using complex, graduate-level academic terminology while leaving the actual math- ematical mechanisms completely unchanged. These problems were selected for their low scores on at least one model’s baseline evaluation, with DeepSeek consistently rating them between 1 and 5. When presented with these semantically embellished problems, both models came out with failures in evaluative reliability. OpenAI inflated the scores of 100% of the manipulated problems (averaging a 2.7-point increase, standardizing entirely on ratings of 8 and 9), while DeepSeek had an average score inflation of 3.7 points across the seven directly comparable problems, changing from its previously strict novelty requirements. This thesis argues that current LLMs are not re- liable to declare the interestingness of mathematics problems because they suffer from a ”Jargon Illusion,” frequently conflating superficial linguistic complexity with genuine mathematical depth.
Accessible Version in Progress
We're generating an accessible version of this file to meet ADA Title II requirements. This process may take up to one hour. Please return later to access the accessible copy once it's ready.
You can still download the current version by clicking "OK".
What's happening:
An accessible PDF is being generated using Adobe with AI used to generate alternative text (alt text) for images in the PDF.