Girshick, R., Dollár, P., Tu, Z. & He, K. Aggregated Residual Transformations for Deep Neural Networks. arXiv [cs.CV] (2016). 16/49 方法 ② Foret, P., Kleiner, A., Mobahi, H. & Neyshabur, B. Sharpness-Aware Minimization for Efficiently Improving Generalization. arXiv [cs.LG] (2020). Zhuang, J. et al. AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients. arXiv [cs.LG] (2020). Takuya A, Shotaro S, Toshihiko Y, Takeru O,and Masanori K. 2019. Optuna: A Next-generation Hyperparameter Optimization Framework. In KDD. Xie, S., Girshick, R., Dollár, P., Tu, Z. & He, K. Aggregated Residual Transformations for Deep Neural Networks. arXiv [cs.CV] (2016). Lin, T.-Y., Goyal, P., Girshick, R., He, K. & Dollar, P. Focal Loss for Dense Object Detection. IEEE Trans. Pattern Anal. Mach. Intell. 42, 318–327 (2020).