大規模な分散システムの運用に関する論文に気づく. • Microsoft [Chen,2020], Uber Technologies [Lee,2024] • 論文の内容を実システムの運用に役立てられないか? 3 学術論文の例 [Koyama, 2026] Chen, Junjie, et al. "How incidental are the incidents? characterizing and prioritizing incidents for large-scale online service systems." Proceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering. 2020. Lee, I-Ting Angelina, et al. "The tale of errors in microservices." Proceedings of the ACM on Measurement and Analysis of Computing Systems 8.3 (2024): 1-36. Koyama, Tomoyuki, Takayuki Kushida, and Soichiro Ikuno. "Root Cause Analysis for Middleware Issues by Kubernetes Resource Events." 2026 18th International Conference on Knowledge and Smart Technology (KST). IEEE, 2026.
• Baiduでの過去4年分のハードウェア故障の 作業チケットを分析 [Wang, 2017] • HDDの故障は81.84%で最多 • SSDの故障は0.31% • HDDのプラッタ(皿)は物理的に回転している ため,摩耗故障が起こりうる (cf. バスタブ曲線) 9 Wang, Guosai, Lifei Zhang, and Wei Xu. "What can we learn from four years of data center hardware failures?." 2017 47th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). IEEE, 2017. Sankar, Sriram, et al. "Datacenter scale evaluation of the impact of temperature on hard disk drive failures." ACM Transactions on Storage (TOS) 9.2 (2013): 1-24. データセンタでのハードウェア故障の上位5件 Baidu [Wang, 2017] Microsoft [Sankar, 2013] 1位 HDD 81.84% HDD 71.1% 2位 その他 10.20% その他 7.7% 3位 メモリ 3.06% 置き換えた マシン 5.6% 4位 電源 1.74% メモリ 5.2% 5位 RAID カード 1.23% 電源 4.0% >
[Jiang,2008] →NetApp • 分散ストレージシステムの可用性の分析 [Ford,2010] → Google 10 Barroso, L. A., et al. “The datacenter as a computer: Designing warehouse-scale machines.” Springer Nature. 2019. Jiang, Weihang, et al. "Are disks the dominant contributor for storage failures? A comprehensive study of storage subsystem failure characteristics." ACM Transactions on Storage (TOS) 4.3 (2008): 1-25. Ford, Daniel, et al. "Availability in globally distributed storage systems." 9th USENIX Symposium on Operating Systems Design and Implementation (OSDI 10). 2010.
1カ所に連鎖(深さ=2)する場合が最多 • 全体の88.8%が他の箇所に連鎖 • Uber Technologies社でマイクロサービス間の RPCのエラーを解析 [Lee,2024] • 80%以上のエラーが3カ所以下に連鎖 • 最も深いケース: 深さ=26 • 連鎖障害のほとんどは連鎖数が3以下 15 Li, Xiaoyun, et al. "Going through the life cycle of faults in clouds: Guidelines on fault handling." 2022 IEEE 33rd International Symposium on Software Reliability Engineering (ISSRE). 2022. Lee, I-Ting Angelina, et al. "The tale of errors in microservices." Proceedings of the ACM on Measurement and Analysis of Computing Systems 8.3 (2024): 1-36. MS MS MS 1 2 深さ=3の例 最初のエラー 3 深さ サーベイ [Li,2022] Uber [Lee,2024] 1 11.9% 33.49% 2 52.6% 37.48% 3 27.1% 11.24% 4 8.2% 4.41% 5 0.3% 8.43% 6≦ 0% 4.94%
ストレージやネットワークの障害がアプリケーション に伝搬している • 典型的なパターン: Li, Xiaoyun, et al. "Going through the life cycle of faults in clouds: Guidelines on fault handling." 2022 IEEE 33rd International Symposium on Software Reliability Engineering (ISSRE). 2022. 16 From To 論文[Li,2022]のTable 4をもとに作成 右図の上位5件 From To 1位 ストレージ アプリケーション 2位 ネットワーク アプリケーション 3位 ミドルウェア アプリケーション 4位 バックエンド アプリケーション 5位 アプリケーション アプリケーション 下位層(インフラ) 上位層(アプリ)
ミドルウェア アプリケーション OS ハードウェア ミドルウェア アプリケーション OS ハードウェア ネットワーク ソフトウェア インフラ マシン1 マシン2 バグの由来[Hopper,’47] Hopper, Grace. "Log Book With Computer Bug." National Museum of American History (1947).
Haopeng, et al. "What bugs cause production cloud incidents?." Proceedings of the Workshop on Hot Topics in Operating Systems. 2019. * In Incident Response, It’s the People Who Make all the Difference | PagerDuty https:// www.pagerduty.com/eng/in-incident-response-its-the-people-who-make-all-the-difference/ Microsoft Azureの本番環境で6ヶ月間に発生した112件のインシデントを ソフトウェアのバグに関して分類 [Liu,2019] コンポーネントの故障 コンポーネントでの障害の検知や対処が不正確または欠落 31% データ形式のバグ 異なるソフトウェアコンポーネント間でデータ形式が衝突 21% タイミングのバグ 永続データやキャッシュデータがエラーなく破損や不整合 13% 定数のバグ ハードコードされた定数の誤った設定(例: タイポ) 7% その他 リソースリーク,セマンティクスのバグ 28%
故障箇所 • ソフトウェアコンポーネントGでFの故障を 検知 22 故障検知のロジックの不足 GがFの故障の可能性に関して 故障検知のロジックをもって いない. G F エラーのシグナルやログが故障で消失 Fの故障でシグナルやログ 自体が消失する可能性に 気づいていない. FからGへの伝送経路で シグナルやログが消失する 可能性に気づいていない. 監視 ログファイル G F Wait G F Liu, Haopeng, et al. "What bugs cause production cloud incidents?." Proceedings of the Workshop on Hot Topics in Operating Systems. 2019.
ArgoCDでクラスターの構成を変更 • 必要な括弧 {{ }} が欠けていた • 新しい構成に有効なnamespaceがないため,世界中のAZとリージョンのすべて のnamespaceにある478のサービスの一括削除が開始された 26 - cell: {{ default $cluster.name $cluster.unique_name }} + cell: $cluster.name Barroso, L. A., et al. “The datacenter as a computer: Designing warehouse-scale machines.” Springer Nature. 2019. Yin, Zuoning, et al. "An empirical study on configuration errors in commercial and open source systems." Proceedings of the Twenty-Third ACM Symposium on Operating Systems Principles. 2011. How a couple of characters brought down our site | by Skyscanner Engineering | Medium https://medium.com/@SkyscannerEng/how-a-couple-of-characters-brought-down-our- site-356ccaf1fbc3
Server + PHP) コンフィグミスあり log_output="Table" log=query.log コンフィグミスあり 値の不整合(MySQL) ログをファイルに書き出そうとするが ログの書き出し先はDBのテーブル ▶コンフィグだけでは何を意図しているか不明 モジュールが見つからずにSEGV →recode.soを先に読み込むべき Yin, Zuoning, et al. "An empirical study on configuration errors in commercial and open source systems." Proceedings of the Twenty-Third ACM Symposium on Operating Systems Principles. 2011.
via multimodal anomaly detection for online service systems." Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 2021. • Ghosh, Supriyo, et al. "How to fight production incidents? an empirical study on a large-scale cloud service." Proceedings of the 13th Symposium on Cloud Computing. 2022. • Li, Liqun, et al. "Fighting the fog of war: Automated incident detection for cloud systems." 2021 USENIX Annual Technical Conference (USENIX ATC 21). 2021. • Shen, Junxian, et al. "Network-centric distributed tracing with deepflow: Troubleshooting your microservices in zero code." Proceedings of the ACM SIGCOMM 2023 Conference. 2023. • Xie, Zhe, et al. "Microservice root cause analysis with limited observability through intervention recognition in the latent space." Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining. 2024. 33
al. "An empirical study on configuration errors in commercial and open source systems." Proceedings of the Twenty- Third ACM Symposium on Operating Systems Principles. 2011.