(2025). On the mechanism of reasoning pattern selection in reinforcement learning for language models. arXiv preprint arXiv:2506.04695. • Fan, Y., Du, Y., Ramchandran, K., & Lee, K. (2024). Looped transformers for length generalization. The Twelfth International Conference on Learning Representations. • Horner, V., & Whiten, A. (2005). Causal knowledge and imitation/emulation switching in chimpanzees (Pan troglodytes) and children (Homo sapiens). Animal Cognition, 8(3), 164–181. https://doi.org/10.1007/s10071-0040239-6 • Huang, Y., Cheng, X., & Liang, Y. (2025). Transformers provably learn chain-of-thought reasoning with length generalization. International Conference on Learning Representations. • Wen, X., Liu, Z., Zheng, S., Xu, Z., Ye, S., Wu, Z., Liang, X., Wang, Y., Li, J., Miao, Z., Bian, J., & Yang, M. (2025). Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base LLMs. arXiv preprint arXiv:2506.14245. • Yuan, H., Xu, Z., Wang, H., Yi, X., Gao, J., Zhang, X.-P., Wang, Y., Yu, C., & Wu, Y. (2026). Verifiable process rewards for agentic reasoning. arXiv preprint arXiv:2605.10325. • Zhang, Z., Chen, Z., Li, M., Tu, Z., & Li, X. (2025). RLVMR: Reinforcement learning with verifiable meta-reasoning rewards for robust long-horizon agents. arXiv preprint arXiv:2507.22844.