Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse Domains. Proceedings of ACL 2026 (Volume 1: Long Papers), pp. 3872–3892. [2] Dian Yu et al. 2021. Self-Teaching Machines to Read and Comprehend with Large-Scale Multi-Subject Question-Answering Data. Findings of EMNLP 2021, pp. 56–68. [3] Qwen Team. 2024. Qwen2.5: A Party of Foundation Models. Qwen Blog. [4] Ronald J. Williams. 1992. Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning. Machine Learning, 8:229–256. [5] Arash Ahmadian et al. 2024. Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMs. ACL 2024, pp. 12248–12267. [6] Zhihong Shao et al. 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300. [7] Weizhe Yuan et al. 2025. NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions. NeurIPS 2025, Datasets and Benchmarks Track. [8] Xiang Yue et al. 2024. MAmmoTH2: Scaling Instructions from the Web. NeurIPS 2024. 17