Researchers have discovered that large language models are surprisingly good at solving complex math problems, but they're not necessarily understanding the underlying math concepts. In fact, these models often rely on shortcuts or tricks to get the right answers, rather than truly grasping the mathematical principles involved. To better understand how these models work and improve their performance, scientists have developed a new framework that evaluates their ability to reason mathematically in four key areas: discovering new ideas, generating solutions, digesting complex concepts, and executing calculations. By identifying which of these areas is weakest, researchers can create more effective training methods that help the models build a stronger foundation in math.