Recent LLMs are approaching and surpassing world-class mathematical ability. Identifying the correct answer given problems and traces that the average human cannot solve becomes a central challenge in achieving artificial superintelligence. Solving a problem with LLMs involves two steps: generating candidate solutions and verifying them. Through our AIMO 3 effort, we found that the relative difficulty of these steps is task-dependent and that competition mathematics is "easy to guess, hard to verify," in contrast to tasks like integer factorization or multi-hop QA benchmarks, where a single trace can be checked easily. This talk introduces existing approaches and discusses challenges for future AI-for-Math research. It also examines competition dynamics — why optimizing the mean score loses to optimizing the upper confidence bound — and the practical challenges of training math models, concluding with our publicly released model and SFT dataset.