Upgrade to Pro — share decks privately, control downloads, hide ads and more …

Fair by design: orchestrating background jobs i...

Sponsored · SiteGround - Reliable hosting with speed, security, and support you can count on.

Fair by design: orchestrating background jobs in Ruby

Are you treating your users fairly? They could be stuck in the queue while a greedy user monopolizes resources. And you might not even know it! In this post, you’ll see if it’s time for you to take background job prioritization seriously—and how to make it fair for all users.

Avatar for Alexander Baygeldin

Alexander Baygeldin

July 16, 2026

More Decks by Alexander Baygeldin

Other Decks in Technology

Transcript

  1. 🧑🍳 🧑🍳 🧑🍳 🍕 🍕 Latency = the time between

    placing the order and starting to prepare it
  2. 🧑🍳 🧑🍳 🧑🍳 🍕 🍕 🧑🍳 🧑🍳 🧢 🧢 🧑🍳

    🧢 🧑🍳 🧢 🧑🍳 🧢 🧑🍳 🧢 🧑🍳 🧢 🧑🍳 🧢 🧑🍳 🧢 🧑🍳 🧢 part-time chefs
  3. 🧑🍳 🧑🍳 🧑🍳 🍕 🍕 🧑🍳 🧑🍳 🧢 🧢 🧑🍳

    🧢 🧑🍳 🧢 🧑🍳 🧢 🧑🍳 🧢 🧑🍳 🧢 🧑🍳 🧢 🧑🍳 🧢 🧑🍳 🧢 part-time chefs At some point improving Quality of Service at more operational cost leads to diminishing returns
  4. G = n ∑ i=1 n ∑ j=1 xi −

    xj 2n2 ¯ x = n ∑ i=1 n ∑ j=1 xi − xj 2n n ∑ i=1 xi 𝒥 (x1 , x2 , …, xn ) = (∑n i=1 xi )2 n ⋅ ∑n i=1 xi 2 = x2 x2 = 1 1 + ̂ cv 2 Jain's index? Gini index? Some other guy's index?
  5. How do we know if we're not? ☑ FIFO queue

    ☑ High latency ☑ Greedy users Are we OK?
  6. 🧑🍳 🧑🍳 🧑🍳 LMOVE queue:tenant_1 queue:sq|<worker ID>|tenant_1 RIGHT LEFT LMOVE

    queue:tenant_2 queue:sq|<worker ID>|tenant_2 RIGHT LEFT ... LMOVE queue:tenant_N queue:sq|<worker ID>|tenant_N RIGHT LEFT worker's "in-progress" queues per-tenant queues
  7. 🧑🍳 🧑🍳 🧑🍳 Solid Queue SELECT job_id FROM solid_queue_ready_executions WHERE

    queue_name = 'tenant_N' ORDER BY priority ASC, job_id ASC LIMIT ? FOR UPDATE SKIP LOCKED
  8. 🧑🍳 🧑🍳 🧑🍳 GoodJob 👍 WITH rows AS MATERIALIZED (

    SELECT id, active_job_id FROM good_jobs WHERE queue_name = 'tenant_N' AND (<more filters>) ORDER BY priority DESC NULLS LAST, created_at ASC LIMIT ? ) SELECT id FROM rows WHERE pg_try_advisory_lock(<lock hash based on active_job_id>) LIMIT 1
  9. (with what we have) 1. Shuf fl e-sharding 2. Interruptible

    iteration 3. Throttling 4. Per-tenant queues Let's fix this!
  10. 🍕 🍕 🍕 🍕 🍕 🍕 Alex 🍕 Amy 🍕

    Joe 🍕 Sam 🍕 Sam 🍕 Amy 🍕 A to I J to R S to Z 🧑🍳 🧑🍳 🧑🍳
  11. 🧑🍳 🧑🍳 🧑🍳 Alex 🍕 Amy 🍕 Joe 🍕 Sam

    🍕 Sam 🍕 Amy 🍕 A to I J to R S to Z
  12. 🧑🍳 🐷 Pentagon 🍕 🧑🍳 🧑🍳 Amy 🍕 Amy 🍕

    Sam 🍕 🐷 Pentagon 🍕 A to I J to R S to Z
  13. 🧑🍳 🐷 Pentagon 🍕 🧑🍳 🧑🍳 Amy 🍕 Amy 🍕

    Sam 🍕 🐷 Pentagon 🍕 A to I J to R S to Z
  14. 🧑🍳 🐷 Pentagon 🍕 🧑🍳 🧑🍳 A to I J

    to R S to Z Amy 🍕 🐷 Pentagon 🍕 💤 Joe 🍕
  15. 🧑🍳 🐷 Pentagon 🍕 🧑🍳 🧑🍳 A to I J

    to R S to Z Amy 🍕 🐷 Pentagon 🍕 Joe 🍕 J to R
  16. 🧑🍳 🐷 Pentagon 🍕 🧑🍳 🧑🍳 A to I J

    to R S to Z Amy 🍕 🐷 Pentagon 🍕 Joe 🍕 J to R Sam 🍕
  17. 🧑🍳 🐷 Pentagon 🍕 🧑🍳 🧑🍳 A to I J

    to R S to Z Amy 🍕 🐷 Pentagon 🍕 Joe 🍕 J to R Sam 🍕 A to I S to Z
  18. Shuffle-sharding TL;DR: good when you have enough workload to fi

    ll all shards. No hogging? No—but it's less likely by a factor of <shard count>. Perfectly fair? No—but it affects fewer people when it's not fair. Full resource utilization? Usually, yes—unless there are too many shards. Does it scale? Yes—in fact, it gets more effective.
  19. Interruptible iteration No hogging? No—it could happen if one tenant

    enqueues too many batches simultaneously. Perfectly fair? No—it stops being fair when someone hogs the queue. Full resource utilization? Usually, yes—except when there are less than <worker count> jobs in the queue. Does it scale? Yes—using cursors to track progress is cheap. TL;DR: good when the workload comes in large batches.
  20. 🪣 🚿 🕳 💧 💧 💧 💧 💧 💧 💧

    💧 💧 over time, the bucket leaks allowing further orders bucket has a limited capacity each new order fi lls the bucket
  21. 🪣 🚿 🕳 💧 💧 💧 💧 💧 💧 💧

    💧 💦 💦 uh-oh, over fl ow! bucket has a limited capacity each new order fi lls the bucket over time, the bucket leaks allowing further orders
  22. 🧑🍳 🧑🍳 🧑🍳 🍕 🍕 🍕 🍕 🎲 default queue

    (80% chance to be chosen) throttled queue (20% chance to be chosen)
  23. 🧑🍳 🧑🍳 🧑🍳 🍕 🍕 🍕 🍕 🎲 default queue

    (80% chance to be chosen) throttled queue (20% chance to be chosen) 🍕
  24. 🧑🍳 🧑🍳 🧑🍳 🍕 🍕 🎲 default queue (80% chance

    to be chosen) throttled queue (20% chance to be chosen) 🍕
  25. Throttling No hogging? Yes—everyone will get at least a little

    work done. Perfectly fair? No—especially if the workload is bursty by nature. Full resource utilization? Yes—with weighted queues support. Does it scale? Yes—especially with leaky buckets. TL;DR: good when the workload is well distributed over time.
  26. 🍕 🍕 🍕 🎲 🍕 🍕 🍕 🍕 🍕 🧑🍳

    🧑🍳 🧑🍳 📒 🍕 🍕 🍕 🧙
  27. Per-tenant queues + custom scheduler No hogging? Yes—it's basically communism.

    Perfectly fair? Yes—as fair as you care to make it. Full resource utilization? Yes—with zero changes to the underlying infra. Does it scale? 🤔
  28. Per-tenant queues + custom scheduler No hogging? Yes—it's basically communism.

    Perfectly fair? Yes—as fair as you care to make it. Full resource utilization? Yes—with zero changes to the underlying infra. Does it scale? Usually, yes—unless the scheduler is slower than workers. TL;DR: good when the workload is heavy enough to forget about the bottleneck.
  29. ✨ AI work fl ows ⬇ Data imports 📊 Report

    generation 🖼 Media processing (where fairness matters) Examples! ` ... and more!
  30. ✨ AI work fl ows ⬇ Data imports 📊 Report

    generation What limits throughput?
  31. ✨ AI work fl ows ⬇ Data imports 📊 Report

    generation What limits throughput? I/O wait
  32. ✨ AI work fl ows ⬇ Data imports 📊 Report

    generation I/O wait API rate limits Database GPU* * if self-hosted What limits throughput?
  33. So, which one? Shuf fl e- sharding Interruptible iteration Throttling

    Per-tenant queues No hogging? ❌ ❌ ✅ ✅ Perfectly fair? ❌ ❌ ❌ ✅ Full resource utilization? 🟧 🟧 ✅ ✅ Does it scale? ✅ ✅ ✅ 🟧