Upgrade to Pro — share decks privately, control downloads, hide ads and more …

Rook: Intro and Deep Dive with Ceph

Rook: Intro and Deep Dive with Ceph

Avatar for Satoru Takeuchi

Satoru Takeuchi PRO

August 01, 2026

More Decks by Satoru Takeuchi

Other Decks in Technology

Transcript

  1. #KubeCon #CloudNativeCon Rook: Intro and Deep Dive with Ceph Dan

    van der Ster & Deepika Upadhyay, Clyso Satoru Takeuchi, Cybozu, Inc. 1
  2. Agenda • Introduction to Rook • • • Introduction to

    Ceph Advanced Features of Rook Case Studies 2
  3. Rook • • • An open source K8s operator to

    manage Ceph storage Administrators can deploy and upgrade Ceph cluster via CRs Users can consume Ceph by PVC and OBC 4
  4. Ceph • • • • • • All-in-one open source

    distributed storage platform ◦ RBD: Network block storage ◦ CephFS: Distributed filesystem ◦ RGW: S3 compatible object storage High scalability: e.g. over 2 EiB Cluster in a real-world deployments High durability: Data replication or EC considering failure domains ◦ Automatic self healing High availability: Almost all operations can be done online High performance: supports AI workloads at high throughputs and IOPS Enterprise features: multi-site, business continuity, tiering, etc… 5
  5. Ceph’s Architecture • • • OSD daemons - “Object Store”

    ◦ Manage disks MON daemons - “Monitor” ◦ Manage cluster’s state & config MGR daemons - “Manager” ◦ Cluster orchestration & add-ons Ceph cluster MGR MGR network storage (e.g. RBD) storage pool MON MON OSD OSD … OSD disk disk … disk 6
  6. Rook’s Architecture • • manage Rook operator Rook ceph-csi ◦

    Manage Ceph cluster manage manage ◦ Provision a pod for each Ceph daemon ceph-csi network storage (e.g. RBD) Ceph cluster ◦ A CSI driver for Ceph ◦ Provision storage from Ceph MGR MGR storage pool MON MON OSD OSD … OSD disk disk … disk 7
  7. For Admins: Deploy Ceph Cluster watch kind: CephCluster metadata: name:

    my-cluster spec: storage: storageClassDeviceSets: - count: 100 manage Rook Ceph cluster (100 OSDs) node node node Change as you need disk disk K8s resource 8
  8. For Admins: Provision Storage Ceph cluster kind:CephBlockPool pool for block

    device Kind: StorageClass create watch kind: CephFilesystem create Rook pool for filesystem admin kind: StorageClass pool for object storage kind: CephObjectStore kind: StorageClass K8s resource 9
  9. For Users: Consume Storage kind:CephBlockPool Kind: StorageClass kind: CephFilesystem Kind:

    PersistentVolumeClaim create kind: StorageClass kind: CephObjectStore kind: StorageClass Kind: PersistentVolumeClaim user Kind: ObjectBucketClaim K8s resource 10
  10. Supported Configurations of PVCs Storage type Volume Mode Access Mode

    Volume Expansion, snapshot, and cloning RBD Block, Filesystem RWO, RWOP, ROX ✅ CephFS Filesystem Above modes and RWX ✅ 11
  11. Ceph Project Status - July 2026 • Active Releases: ◦

    ◦ • Tentacle (v20.2.2): ◦ ◦ • Dramatically improves erasure coding performance for faster, more cost-efficient storage. Enhanced gateway support with major upgrades to Samba, NFS, and NVMe-oF. Umbrella (v21) ◦ • The current active Ceph releases are Squid (v19) and Tentacle (v20). Next version Umbrella (v21) is in feature freeze. Even faster EC; S3 object dedup; NFS-Ganesha active-active; managed SMB; … Ceph is finally joining the Linux Foundation: ◦ ◦ Ceph Foundation was created in 2018 - Not for Profit in support of the Ceph Project Recently “Ceph Project” has been created in the LF → improved technical governance!! 13
  12. Ceph Project Status - July 2026 CRIMSON MCLOCK/QOS MULTISITE ORCHESTRATOR

    / ROOK DASH/AUTOSCALE/S3 TIERING UPMAP BLUESTORE/MGR/MULTIMDS FULLY AWESOME FS SCRUB ERASURE CODING 14
  13. Ceph Project Status - July 2026 WE ARE HERE Pacific

    Mar 2021 Quincy Mar 2022 16.2.z • • • Reef Jun 2023 17.2.z Squid Sep 2024 18.2.z Tentacle Nov 2025 19.2.z Stable, named release every 12 months Backports for 2 releases ◦ 18.2.8 was the final Reef release Upgrade up to 2 releases at a time ◦ Pacific → Reef, Quincy → Squid, Reef → Tentacle → Umbrella Umbrella ~Sep 2026 20.2.z 15 15
  14. Remote Replication Storage type Replication Feature Custom Resource RBD RBD

    mirroring CephRBDMirror CephFS CephFS mirroring CephFilesystemMirror RGW RGW multisite CephObjectRealm K8s resource 21
  15. NVMe-oF • Export RBD device as NVMe device ◦ Consume

    RBD on nodes that don’t have Ceph client ◦ Far faster than iSCSI (maintenance mode) ◦ Consume devices from internal/external K8s Clusters and non-K8s envs watch create kind:CephBlockPool Rook Create kind:CephNVMeOFGateway admin Kind: StorageClass NVMe-oF gateway K8s resource 22
  16. RBD QoS • Set I/O limits of RBD devices via

    VolumeAttributeClass CRs ◦ Max read/write IOPS ◦ Max read/write Bps create Kind: VolumeAttributeClass create Kind: PersistentVolumeClaim admin user K8s resource 23
  17. COSI (upcoming) • • K8s official interface for object storage

    Stability ◦ Implementing new v1alpha2 API ◦ Rook supports v1alpha1 but the interface will change drastically create kind:BucketClass admin kind:BucketClaim create user K8s resource 24
  18. Performance Envelope • • Workload: 2M objects (4 KB each),

    20 K6 clients, single bucket. Method: Ramping load over 30 minutes until p90 latency exceeded 500 ms. ~90K GET/s saturation begins ~170K GET/s breaking point
  19. Day 1: Infrastructure as a Code • GitOps • Base

    Blueprint as Helm Chart • Region Overlay: CephBlockPools, RGW placement rules.
  20. Conclusion • • • Ceph is an open source distributed

    storage Rook is a Ceph orchestration on K8s Rook makes managing/using Ceph a lot easier 34
  21. Philosophy • • Support latest Ceph and K8s Make Ceph

    the best storage platform for K8s 35
  22. Stability • • • CNCF graduated project Marked as stable

    8 years ago Many users running Rook in production 36