GeoEval : benchmark for evaluating LLMs and Multi-Modal Models on geometry problem-solving
Zhang, Jiaxin and Li, Zhongzhi and Zhang, Mingliang and Yin, Fei and Liu, Chenglin and Moshfeghi, Yashar (2024) GeoEval : benchmark for evaluating LLMs and Multi-Modal Models on geometry problem-solving. Other. arXiv, Ithaca, NY. (https://doi.org/10.48550/arXiv.2402.10104)
Preview |
Text.
Filename: Zhang-etal-arXiv-2024-GeoEval-benchmark-for-evaluating-LLMs-and-Multi-Modal-Models.pdf
Final Published Version License: Download (5MB)| Preview |
Abstract
Recent advancements in Large Language Models (LLMs) and Multi-Modal Models (MMs) have demonstrated their remarkable capabilities in problem-solving. Yet, their proficiency in tackling geometry math problems, which necessitates an integrated understanding of both textual and visual information, has not been thoroughly evaluated. To address this gap, we introduce the GeoEval benchmark, a comprehensive collection that includes a main subset of 2000 problems, a 750 problem subset focusing on backward reasoning, an augmented subset of 2000 problems, and a hard subset of 300 problems. This benchmark facilitates a deeper investigation into the performance of LLMs and MMs on solving geometry math problems. Our evaluation of ten LLMs and MMs across these varied subsets reveals that the WizardMath model excels, achieving a 55.67\% accuracy rate on the main subset but only a 6.00\% accuracy on the challenging subset. This highlights the critical need for testing models against datasets on which they have not been pre-trained. Additionally, our findings indicate that GPT-series models perform more effectively on problems they have rephrased, suggesting a promising method for enhancing model capabilities.
ORCID iDs
Zhang, Jiaxin ORCID: https://orcid.org/0000-0001-7355-7975, Li, Zhongzhi, Zhang, Mingliang, Yin, Fei, Liu, Chenglin and Moshfeghi, Yashar ORCID: https://orcid.org/0000-0003-4186-1088;-
-
Item type: Monograph(Other) ID code: 89011 Dates: DateEvent15 February 2024PublishedSubjects: Science > Mathematics > Electronic computers. Computer science > Other topics, A-Z > Human-computer interaction Department: Faculty of Science > Computer and Information Sciences Depositing user: Pure Administrator Date deposited: 29 Apr 2024 09:13 Last modified: 11 Nov 2024 16:08 URI: https://strathprints.strath.ac.uk/id/eprint/89011