<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd"><dc:title>Assessing the Reliability of Large Language Models in Evaluating Mathematical Interestingness</dc:title><dc:creator>Arora, Varnikaa </dc:creator><dc:subject>Large Language Models</dc:subject><dc:subject>Mathematical Interestingness</dc:subject><dc:subject>Automated Mathematical Discovery</dc:subject><dc:subject>Jargon Illusion</dc:subject><dc:subject>Semantic Equivalence</dc:subject><dc:subject>Linguistic Bias</dc:subject><dc:subject>Prompt Sensitivity</dc:subject><dc:subject>Heuristics</dc:subject><dc:subject>Neurosymbolic Architecture</dc:subject><dc:subject>GPT-4o</dc:subject><dc:subject>DeepSeek-V3</dc:subject><dc:subject>Large Language Models</dc:subject><dc:subject>Mathematical Interestingness</dc:subject><dc:subject>Automated Mathematical Discovery</dc:subject><dc:subject>Jargon Illusion</dc:subject><dc:subject>Semantic Equivalence</dc:subject><dc:subject>Linguistic Bias</dc:subject><dc:subject>Prompt Sensitivity</dc:subject><dc:subject>Heuristics</dc:subject><dc:subject>Neurosymbolic Architecture</dc:subject><dc:subject>GPT-4o</dc:subject><dc:subject>DeepSeek-V3</dc:subject><dc:coverage>Computer Science</dc:coverage><dc:relation>B S</dc:relation><dc:description>The question of what makes a mathematical problem interesting is important to both mathe-
matics education and the field of automated mathematical discovery (AMD). As large language
models (LLMs) are increasingly used to evaluate, generate, and prioritize mathematical content,
understanding whether different models apply consistent and coherent criteria for interesting-
ness becomes an important factor. This thesis investigates whether LLMs (OpenAI GPT-4o and
DeepSeek-V3) give comparable ratings of interestingness when evaluating identical mathematical
problems, and tests the structural integrity of their underlying evaluative rubrics.
Using a structured prompting methodology with a reproducible random seed, both models
were applied to 100 problems drawn from the 42 math dataset. Initially, the models agreed exactly
on only 44% of problems, showing substantial divergence on 24% of problems (up to a 5-point
difference on a 10-point scale). Analysis revealed this divergence came from conflicting default
heuristics. DeepSeek applied a strict novelty-and-non-obviousness standard, which resulted in a
strongly bimodal distribution, while OpenAI correlated foundational, historical significance with
interestingness.
However, to determine if these differing heuristics were grounded in actual mathematical com-
prehension, a ”Semantic Equivalence” test was conducted. Ten mathematically trivial problems
were rewritten using complex, graduate-level academic terminology while leaving the actual math-
ematical mechanisms completely unchanged. These problems were selected for their low scores
on at least one model’s baseline evaluation, with DeepSeek consistently rating them between 1
and 5. When presented with these semantically embellished problems, both models came out with
failures in evaluative reliability. OpenAI inflated the scores of 100% of the manipulated problems
(averaging a 2.7-point increase, standardizing entirely on ratings of 8 and 9), while DeepSeek had
an average score inflation of 3.7 points across the seven directly comparable problems, changing
from its previously strict novelty requirements. This thesis argues that current LLMs are not re-
liable to declare the interestingness of mathematics problems because they suffer from a ”Jargon
Illusion,” frequently conflating superficial linguistic complexity with genuine mathematical depth.</dc:description><dc:contributor>Abhinav Verma, Thesis Supervisor</dc:contributor><dc:contributor>Rebecca Jane Passonneau, Thesis Honors Advisor</dc:contributor><dc:rights>open_access</dc:rights><dc:date>2026-04-03T05:27:13Z</dc:date><dc:identifier>https://honors.libraries.psu.edu/catalog/10345vpa5115</dc:identifier></oai_dc:dc>