Upgrade to Pro — share decks privately, control downloads, hide ads and more …

[seminar talk'26] End-to-End Training of Two-St...

[seminar talk'26] End-to-End Training of Two-Stage Decision Systems: Towards Personalized Decisions at Scale

Slides used for the online talk at CSE DSI Machine Learning Seminar Series at University of Minnesota (UMN): https://cse.umn.edu/dsi/events/cse-dsi-machine-learning-seminar-haruka-kiyohara-computer-science-cornell

Avatar for Haruka Kiyohara

Haruka Kiyohara

September 29, 2026

More Decks by Haruka Kiyohara

Other Decks in Research

Transcript

  1. End-to-End Training of Two-Stage Decision Systems: Towards Personalized Decisions at

    Scale Haruka Kiyohara ([email protected]) University of Minnesota, CSE DSI Machine Learning Seminar Series September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 1
  2. About me • 4th year Ph.D. candidate @ Cornell CS

    (advisor: Thorsten Joachims & Sarah Dean) • Working on.. • Recommender Systems (RecSys) • Reinforcement Learning (RL) • Causal Inference (CI) 清原 明加 Haruka Kiyohara [email protected] September 2026 • Published at • ICML, NeurIPS, ICLR, AAAI • KDD, RecSys, WSDM, TheWebConf, CIKM End-to-End Training of Two-Stage Decision Systems @ UMN 2
  3. About me • 4th year Ph.D. candidate @ Cornell CS

    (advisor: Thorsten Joachims & Sarah Dean) • Working on.. • Recommender Systems (RecSys) • Reinforcement Learning (RL) • Causal Inference (CI) 清原 明加 Haruka Kiyohara [email protected] September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 3
  4. About me • 4th year Ph.D. candidate @ Cornell CS

    (advisor: Thorsten Joachims & Sarah Dean) • Working on.. • Recommender Systems (RecSys) • Reinforcement Learning (RL) • Causal Inference (CI) 清原 明加 Haruka Kiyohara [email protected] this Friday @Minneapolis! September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 4
  5. About me • 4th year Ph.D. candidate @ Cornell CS

    (advisor: Thorsten Joachims & Sarah Dean) • Supported by.. • LinkedIn-Cornell Bowers Strategic Partnership (2026-2027) • Funai Overseas Scholarship (2023-2025) • Quad Fellowship (2025-2026) 清原 明加 Haruka Kiyohara [email protected] September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 5
  6. About me • 4th year Ph.D. candidate @ Cornell CS

    (advisor: Thorsten Joachims & Sarah Dean) • Supported by.. • LinkedIn-Cornell Bowers Strategic Partnership (2026-2027) • Funai Overseas Scholarship (2023-2025) • Quad Fellowship (2025-2026) 清原 明加 Haruka Kiyohara [email protected] September 2026 • Misc: • I’m an AI-music composer! suno.com/@sunosuno07 End-to-End Training of Two-Stage Decision Systems @ UMN 6
  7. Machine decision-making systems are everywhere! search recommendation Chatbots/AI assistance (LLMs)

    September 2026 SNS Creatives (GenAI) End-to-End Training of Two-Stage Decision Systems @ UMN 7
  8. Machine decision-making systems are everywhere! search recommendation SNS • We

    need to handle massive volume of items (e.g., movies, webpages, SNS posts) • We need to handle inference at web-latency (~0.02 second) • We’d like to unlock the capability of cutting-edge models (e.g., LLMs, diffusions) Chatbots/AI assistance (LLMs) September 2026 Creatives (GenAI) End-to-End Training of Two-Stage Decision Systems @ UMN 8
  9. Machine decision-making systems are everywhere! search recommendation SNS • We

    need to handle massive volume of items (e.g., movies, webpages, SNS posts) • We need to handle inference at web-latency (~0.02 second) • We’d like to unlock the capability of cutting-edge models (e.g., LLMs, diffusions) Chatbots/AI assistance (LLMs) A good Creatives (GenAI) tradeoff between “inference latency” and “model expressiveness” is the key! September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 9
  10. Two-stage decision-making systems In large-scale decision systems, we often employ

    a two-stage selection. Large-scale recommendation (RecSys) candidate retrieval (i.e., fast screening) September 2026 output generation (e.g., ranker, LLMs) End-to-End Training of Two-Stage Decision Systems @ UMN 10
  11. Two-stage decision-making systems In large-scale decision systems, we often employ

    a two-stage selection. Large-scale recommendation (RecSys) Retrieval-augmented generation (RAG) candidate retrieval (i.e., fast screening) September 2026 output generation (e.g., ranker, LLMs) End-to-End Training of Two-Stage Decision Systems @ UMN 11
  12. Two-stage decision-making systems In orchestrating large-scale decision systems, we also

    use a two-stage decision. early-stage late-stage “Model A execute task 1: translation.” AI orchestration “Model B execute task 2: Q&A.” The “conductor” model provides instructions for a single/multiple AI model(s). September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN “Model C execute task 3: info-graphics.” 12
  13. Research question How can we train two-stage decision process (especially

    its early stage) end-to-end using user feedback (e.g., clicks, purchases)? Large-scale recommendation (RecSys) Optimize early-stage using click signals! September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 13
  14. Topics we will discuss today 1. Candidate Retrieval for Large-Scale

    Decision Systems (~20min) Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking. [KCEKNPRDJW, ICML2026] 2. Orchestration for Large-Scale Decision Systems (~10min) An Off-Policy Learning Approach for Steering Sentence Generation towards Personalization. [KCSJ, RecSys2025] 3. Open Challenges in Two-Stage Decisions (~10min) Intrinsic and Extrinsic Diversity-aware Candidate Retrieval in Two-stage Decisions. [on-going work] Fast and Scalable Algorithms for Adaptable Recommendations. [on-going work] Policy Design for Two-sided Platforms with Participation Dynamics. [KYD, ICML2025] September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 14
  15. Candidate Retrieval for large-scale decisions Credit-assigned Policy Gradient for Early

    Stage Retrieval in Two-stage Ranking. [KCEKNPRDJW, ICML2026] (Internship work at Meta) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 15
  16. Two-stage decision-making systems In large-scale decision systems, we often employ

    a two-stage selection. Large-scale recommendation (RecSys) Retrieval-augmented generation (RAG) candidate retrieval (i.e., fast screening) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 16
  17. Two-stage decision-making systems In large-scale decision systems, we often employ

    a two-stage selection. less research more research Large-scale recommendation (RecSys) How early stage contributes to the outcome? September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 17
  18. Two-stage decision-making systems In large-scale decision systems, we often employ

    a two-stage selection. less research more research Large-scale recommendation (RecSys) How early stage contributes to the outcome? How can we improve the overall decision process by appropriately “assign credits” for the early stage retrieval? September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 18
  19. Overall decision process context candidate set reward ranking where the

    (two stage) policy is defined as: optimize September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN fixed 19
  20. Our goal: Maximize the user feedback We aim to maximize

    the user’s reward signal, such as clicks or purchases. : position weight September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 20
  21. Objective maximization via policy gradient In reinforcement learning (RL), we

    often derive the policy gradient of the objectives. Reinforcement Learning (RL) Supervised Learning (SL) gradient ascent gradient descent September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 21
  22. Objective maximization via policy gradient In reinforcement learning (RL), we

    often derive the policy gradient of the objectives. differentiate with 𝝅 Reinforcement Learning (RL) Supervised Learning (SL) gradient ascent gradient descent September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 22
  23. Target policy gradient of interest Now, consider the gradient of

    the joint policy (ESR + LSR) for position 𝑙’s reward. ESR September 2026 LSR End-to-End Training of Two-Stage Decision Systems @ UMN This specific position 𝑙. 23
  24. Target policy gradient of interest Now, consider the gradient of

    the joint policy (ESR + LSR) for position 𝑙’s reward. ESR LSR This specific position 𝑙. Why? – because we can calculate global policy gradient for the entire ranking by the linear sum of each position gradient. (when there is no item-item interaction) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 24
  25. Baseline: “vanilla” policy gradient (V-PG) Now, consider the gradient of

    the joint policy (ESR + LSR) for position 𝑙’s reward. (marginal) prob. of having action 𝑎𝑙 at position 𝑙 September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 25
  26. Baseline: “vanilla” policy gradient (V-PG) Now, consider the gradient of

    the joint policy (ESR + LSR) for position 𝑙’s reward. Rewriting the prob using two stage policies.. ESR September 2026 LSR End-to-End Training of Two-Stage Decision Systems @ UMN 26
  27. Baseline: “vanilla” policy gradient (V-PG) Now, consider the gradient of

    the joint policy (ESR + LSR) for position 𝑙’s reward. Rewriting the prob using two stage policies.. Propagating the gradient to the ESR’s candidate selection prob. [Ma+,20]. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 27
  28. V-PG looks like a promising approach but.. V-PG does not

    scale when the candidate set size (𝐾) becomes large, due to September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 28
  29. V-PG looks like a promising approach but.. V-PG does not

    scale when the candidate set size (𝐾) becomes large, due to → Even when 𝐴 = 10, the total number of combination (|𝐴|𝐾) becomes exponentially large.. ! (In practice, we may have 𝐴 = 1,000,000 or even more!) # of combinations • Variance issue • Action space scale up expotentially ≈ 𝑂(|𝐴|𝐾 ) • In contrast, we can only sample results with a single candidate set. 10 x 9 x 8 x .. candidate set size (𝐾) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 29
  30. V-PG looks like a promising approach but.. V-PG does not

    scale when the candidate set size (𝐾) becomes large, due to • Variance issue • Action space scale up expotentially ≈ 𝑂(|𝐴|𝐾 ) • In contrast, we can only sample results with a single candidate set. We want to avoid the dependence on candidate set September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN here! 30
  31. Proposal: Credit-Assigned Policy Gradient (CA-PG) We consider a marginalized distribution

    for reducing the variance: September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 31
  32. Proposal: Credit-Assigned Policy Gradient (CA-PG) We consider a marginalized distribution

    for reducing the variance: What is this marginal probability? (marginal) prob. of having action 𝑎𝑙 (marginal) prob. of selecting action 𝑎𝑙 included in top-K (in some candidate) given the fact that 𝑎𝑙 is in (some) top-K September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 32
  33. Proposal: Credit-Assigned Policy Gradient (CA-PG) We consider a marginalized distribution

    for reducing the variance: What is this marginal probability? item pool [A], [B], [C], [D] September 2026 V-PG (baseline) workflow (enumerated) candidate set ({𝐴𝐾}) target action (𝑎𝑙) ESR (marginal) prob. of having action 𝑎𝑙 (marginal)LSR prob. of selecting action 𝑎𝑙 [A,top-K B, C], [A] top-K included in (in some candidate) given the fact that 𝑎𝑙 is in (some) End-to-End Training of Two-Stage Decision Systems @ UMN 33
  34. Proposal: Credit-Assigned Policy Gradient (CA-PG) We consider a marginalized distribution

    for reducing the variance: What is this marginal probability? where item pool [A], [B], [C], [D] September 2026 V-PG (baseline) workflow (enumerated) candidate set ({𝐴𝐾}) target action (𝑎𝑙) ESR (marginal) prob. of having action 𝑎𝑙 (marginal)LSR prob. of selecting action 𝑎𝑙 [A,top-K B, C], [A] top-K included in (in some V-PG;candidate) 𝝅𝐄𝐒𝐑(𝑨𝑲|𝒙)given the fact that 𝑎𝑙 is in (some) End-to-End Training of Two-Stage Decision Systems @ UMN 34
  35. Proposal: Credit-Assigned Policy Gradient (CA-PG) We consider a marginalized distribution

    for reducing the variance: What is this marginal probability? where item pool [A], [B], [C], [D] September 2026 (enumerated) candidate set ({𝐴𝐾}) target action (𝑎𝑙) ESR (marginal) prob. of having action 𝑎𝑙 (marginal)LSR prob. of selecting action 𝑎𝑙 [A,top-K B, C], [A] top-K included in (in some candidate) given the fact that 𝑎𝑙 is in (some) CA-PG; 𝝅𝐄𝐒𝐑(𝑺𝑲(𝐀)|𝒙) [A, B, D], [A, C, D], (marginal) [B, C, D]prob. of having action 𝑎𝑙 included in top-K (in some candidate) End-to-End Training of Two-Stage Decision Systems @ UMN 35
  36. Proposal: Credit-Assigned Policy Gradient (CA-PG) We consider a marginalized distribution

    for reducing the variance: effective size of action What is this marginal probability? where item pool [A], [B], [C], [D] September 2026 (enumerated) candidate set ({𝐴𝐾}) target action (𝑎𝑙) ESR (marginal) prob. of having action 𝑎𝑙 (marginal)LSR prob. of selecting action 𝑎𝑙 [A,top-K B, C], [A] top-K included in (in some candidate) given the fact that 𝑎𝑙 is in (some) CA-PG; 𝝅𝐄𝐒𝐑(𝑺𝑲(𝐀)|𝒙) [A, B, D], [A, C, D], (marginal) [B, C, D]prob. of having action 𝑎𝑙 included in top-K (in some candidate) End-to-End Training of Two-Stage Decision Systems @ UMN 36
  37. Why is it called “credit-assigned PG”? A. Because of the

    difference in the gradient propargation workflow. Vanilla-PG sampled candidate set “Credit-Assigned”-PG marginal of target action choice CA-PG looks more efficient! September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 37
  38. Key theoretical properties (1/2) • CA-PG reduces the variance of

    V-PG! (By avoiding the dependence on specific candidate set) • However, CA-PG is a biased policy gradient. (As it “marginalized” and ignored part of the gradient) (V-PG) = (CA-PG) + (𝐴𝐾 dependent gradient) z CA-PG ignores this term via marginalization September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 39
  39. Key theoretical properties (2/2) • When late-stage policy (LSR) is

    accurate, CA-PG can learn a good ESR! (as long as LSR preserves the policy-weighted reward ratio aligned with the reward ratio) ESR LSR implicit (reward-weighted) teacher distillation (i.e., uses the LSR action choice prob. as a part of reward) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 40
  40. Key theoretical properties (2/2) • When late-stage policy (LSR) is

    accurate, CA-PG can learn a good ESR! (as long as LSR preserves the policy-weighted reward ratio aligned with the reward ratio) expected reward ratio where policy-weighted expected reward ratio z = expected prob of the LSR chooses A provided candidate set under the ESR distribution September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 41
  41. Synthetic experiment better We first see the results when model

    is (almost) well-specified. With the increased candidate set sizes (𝑲), CA-PG shows faster convergence and better stability! September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 42
  42. Real-data experiment on KuaiRec [Gao+,22] better We test with a

    larger size of candidate set (𝐾) in {50, 100, 200}, where 𝐴 = 1000. Observed a result similar to the synthetic setting, CA-PG-SwR converges faster when 𝑲 is large. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 43
  43. Takeaways (Part 1) • We studied how to improve early-stage

    retrieval of two-stage decisions. • To key challenge was high variance and credit-assignment issues of vanilla PG. • We proposed credit-assigned PG, which uses LSR prob as a reward signal. CA-PG improves data-efficiency under a mild condition about the LSR quality! September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 44
  44. Orchestration for large-scale decisions An Off-Policy Learning Approach for Steering

    Sentence Generation towards Personalization. [KCSJ, RecSys2025] (PhD work at Cornell) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 45
  45. Two-stage decision-making systems In orchestrating large-scale decision systems, we also

    use a two-stage decision. early-stage late-stage “Model A execute task 1: translation.” AI orchestration “Model B execute task 2: Q&A.” The “conductor” model provides instructions for a single/multiple AI model(s). September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN “Model C execute task 3: info-graphics.” 46
  46. Example: personalized sentence generation We aim to personalize short product

    summary for recommendations: “WALL-E (2008)” short summary https://movies.disney.com/wall-e September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 47
  47. Example: personalized sentence generation We aim to personalize short product

    summary for recommendations: “WALL-E (2008)” ・A robot called “WALL-E” and his adventure into space ・Animated films with beautiful picture and pretty charactors ・Science-fiction focuing on environment destruction ・Heart-warming drama about love and companionship ・Re-discovery of earth and humanity in dystopia ・Silent film without explicit quotes September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 48
  48. Example: personalized sentence generation We aim to personalize short product

    summary for recommendations: “WALL-E (2008)” For sci-fi lovers, In the distant future, one little robot sparked a cosmic revolution. ・A robot called “WALL-E” and his adventure into space ・Animated films with beautiful picture and pretty charactors For romance lovers, ・Science-fiction focuing on environment destruction In a lonely world, a small robot discovers the power of connection. ・Heart-warming drama about love and companionship ・Re-discovery of earth and humanity in dystopia ・Silent film without explicit quotes September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN We’d like to personalize the sentence to each user. 49
  49. Overall decision process user, movie (personalized) prompt short slogan reward

    where the sentence generation policy is defined as: optimize September 2026 fixed (frozen LLM) End-to-End Training of Two-Stage Decision Systems @ UMN 50
  50. Goal: Off-Policy Learning (OPL) Our goal is to optimize the

    policy to maximize the total reward: September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 51
  51. Goal: Off-Policy Learning (OPL) Our goal is to optimize the

    policy to maximize the total reward: , using the logged data collected by a logging policy 𝜋0. need to deal with the partial rewards and distribution shift September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 52
  52. Baseline: Importance Sampling (IS) on Actions Naive approach estimates the

    policy gradient (PG) to update the prompt policy. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 53
  53. Baseline: Importance Sampling (IS) on Actions Naive approach estimates the

    policy gradient (PG) to update the prompt policy. estimate from logged data 𝑛: data size apply importance sampling on the prompt policy prob [Swaminathan&Joachims,16] September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 54
  54. Baseline: Importance Sampling (IS) on Actions Naive approach estimates the

    policy gradient (PG) to update the prompt policy. estimate from logged data 𝑛: data size apply importance sampling on the prompt policy prob [Swaminathan&Joachims,16] Issues September 2026 • high variance due to rejection sampling on large prompt space • discarding rich information about generated sentence End-to-End Training of Two-Stage Decision Systems @ UMN 55
  55. Proposal: Direct Sentence Off-policy gradient (DSO) We estimate the sentence

    policy gradient using logged data as follows. (estimating (similarity-)marginalized sentence policy gradient) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 56
  56. Proposal: Direct Sentence Off-policy gradient (DSO) We estimate the sentence

    policy gradient using logged data as follows. (estimating (similarity-)marginalized sentence policy gradient) a kernel with bandwidth 𝝉 (1) simulating sentence generation using the current policy (similar to GRPO [Gao+,22]) (2) apply soft weighting based on similarity to logged sentence September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 57
  57. Key theoretical property; flexible bias-variance tradeoff Kernel bandwidth 𝛕 controlls

    the bias-variance tradeoff. (𝜋) (+/- 1) (+/- 5) DSO often achieves better bias-variance tradeoff than action-based IS. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 58
  58. better Synthetic experiment results • DSO particularly works well when

    # of actions and reward noises are large. • DSO is much more data-efficient than the baselines. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 59
  59. better Results on MovieLens data [Harper&Konstan,15] • DSO often performs

    better than other OPL methods. • Especially, DSO is more robust to performance corruption. Note: “policy value” is the improvement observed over the sentences generated without prompt, which we call no-prompt baseline. Experiments results is from 25 different trials. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 60
  60. Takeaways (Part 2) • We discussed orchestration example using personalized

    prompt optimization. • To key point is to consider output distribution rather than prompt distribution. • We achieves better bias-variance tradeoff via GRPO-like data augmentation. We can improve data-efficiency by effectively using the decision structure! September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 61
  61. Open challenges for two-stage decisions Both for Candidate Retrieval and

    Orchestration examples September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 62
  62. Open challenges for Candidate Retrieval (1/2) In the CA-PG example,

    we focused on aligning items based on relevance. Large-scale recommendation (RecSys) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 63
  63. Open challenges for Candidate Retrieval (1/2) In the CA-PG example,

    we focused on aligning items based on relevance. Large-scale recommendation (RecSys) However, in applications like news recommendation, item diversity matters. How can we learn to diversify candidate set from user feedback? Intrinsic and Extrinsic Diversity-aware Candidate Retrieval in Two-stage Decisions. [on-going work] September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 64
  64. Open challenges for Candidate Retrieval (2/2) Another challenge arises when

    there is information gap between two-stages. Large-scale recommendation (RecSys) users’ static features September 2026 search query / user prompt End-to-End Training of Two-Stage Decision Systems @ UMN 65
  65. Open challenges for Candidate Retrieval (2/2) Another challenge arises when

    there is information gap between two-stages. Large-scale recommendation (RecSys) users’ static features search query / user prompt How can we risk-hedge for the late stage decision under the gap? How can we reduce the information gap keeping inference latency? Fast and Scalable Algorithms for Adaptable Recommendations. [on-going work] September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 66
  66. Open challenges for Orchestration (1/2) AI orchestration relies on the

    downstream AI models. updated! temporaly unavailable.. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 67
  67. Open challenges for Orchestration (1/2) AI orchestration relies on the

    downstream AI models. updated! temporaly unavailable.. How can we learn to adapt to model changes train/test-time? How can we make the pipeline robust to runtime uncertainty? September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 68
  68. Open challenges for Orchestration (2/2) We may have some dynamics

    in user participation in the decision-systems. Policy Design for Two-sided Platforms with Participation Dynamics. [KYD, ICML2025] September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 69
  69. Open challenges for Orchestration (2/2) We may have some dynamics

    in user participation in the decision-systems. (While not limited to Orchestration,) AI decisions affect what we have in the future. How can we calibrate decisions for long-term success or safety? Policy Design for Two-sided Platforms with Participation Dynamics. [KYD, ICML2025] September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 70
  70. Happy to take questions! 1. Candidate Retrieval for Large-Scale Decision

    Systems Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking. [KCEKNPRDJW, ICML2026] 2. Orchestration for Large-Scale Decision Systems An Off-Policy Learning Approach for Steering Sentence Generation towards Personalization. [KCSJ, RecSys2025] 3. Open Challenges in Two-Stage Decisions Intrinsic and Extrinsic Diversity-aware Candidate Retrieval in Two-stage Decisions. [on-going work] Fast and Scalable Algorithms for Adaptable Recommendations. [on-going work] Policy Design for Two-sided Platforms with Participation Dynamics. [KYD, ICML2025] September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 71
  71. Acknowledgements + collaborators at + collaborators at Cornell (on-going projects,

    undergrads) Cornell-LinkedIn Strategic Partnership (2023-2025) September 2026 (2025-2026) End-to-End Training of Two-Stage Decision Systems @ UMN (2026-2027) 72
  72. Thank you for listening! Feel free to reach out to

    me: [email protected] September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 73
  73. List of presented papers (my own work) [KCEKNPRDJW, ICML2026] Haruka

    Kiyohara, Mihaela Curmei, Ariel Evnine, Shankar Kalyanaraman, Israel Nir, Ana-Roxana Pop, Nitzan Razin, Sarah Dean, Thorsten Joachims, Udi Weinsberg. Creditassigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking. ICML, 2026. https://arxiv.org/pdf/2605.26385 [KCSJ, RecSys2025] Haruka Kiyohara, Daniel Yiming Cao, Yuta Saito, Thorsten Joachims. An OffPolicy Learning Approach for Steering Sentence Generation towards Personalization. RecSys, 2025. https://dl.acm.org/doi/epdf/10.1145/3705328.3748088 [KYD, ICML2025] Haruka Kiyohara, Fan Yao, Sarah Dean. Policy Design for Two-sided Platforms with Participation Dynamics. ICML, 2025. https://arxiv.org/pdf/2502.01792 Stay tuned for updates! (on-going works) Intrinsic and Extrinsic Diversity-aware Candidate Retrieval in Two-stage Decisions. Fast and Scalable Algorithms for Adaptable Recommendations. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 74
  74. Appendix for the Candidate Retrieval paper Credit-assigned Policy Gradient for

    Early Stage Retrieval in Two-stage Ranking. [KCEKNPRDJW, ICML2026] (Internship work at Meta) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 75
  75. Synthetic experiment (2/4) We first see the results when model

    is (almost) well-specified. When the LSR’s alignment condition is satisfied, CA-PG works well! September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 76
  76. Synthetic experiment (3/4) We first see the results when model

    is (almost) well-specified. Results when using multiple sub-retriever to sample candidate set CA-PG gains benefits from combining multiple sub-retrievers (mixture of experts; MoE) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 77
  77. Synthetic experiment (4/4) Also looking at the computational time of

    each method.. 𝑲 𝑳 𝑲 𝑴 Combined with the SwR approximation, both the computational time does not increase with 𝑲 and 𝑳. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 78
  78. CA-PG: Marginalization for reducing the variance We propose the following

    credit-assigned policy gradient (CA-PG): where (marginal) prob. of having action 𝑎𝑙 (marginal) prob. of selecting action 𝑎𝑙 included in top-K (in some candidate) given the fact that 𝑎𝑙 is in (some) top-K A set of candidate sets 𝐴𝐾 that contains action 𝑎𝑙 Sum of the ESR prob of the candidates September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 79
  79. Theoretical analysis (1/3) 1) What is the relation between the

    two PGs? Now, all the factors that depends on 𝑨𝑲 is ignored! September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 80
  80. Theoretical analysis (1/3) 1) What is the relation between the

    two PGs? By ignoring from which candidate 𝑎𝑙 comes from, CA-PG reduces variance. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 81
  81. Theoretical analysis (2/3) 2) What do V-PG and CA-PG optimizes

    for? LSR’s action choice probability-discounted reward September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 82
  82. Theoretical analysis (2/3) 2) What do V-PG and CA-PG optimizes

    for? CA-PG uses LSR’s choice as a reward signal. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 83
  83. Theoretical analysis (3/3) 3) When CA-PG can learn the accurate

    alignment of actions? Alignment of the LSR’s action choice probability matters September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 84
  84. Theoretical analysis (3/3) 3) When CA-PG can learn the accurate

    alignment of actions? ・Any (oracle) epsilon-greedy and softmax-type policies satisfies the alignment condition of LSR. ・Even when the LSR policy makes mistake in the action alignment, the following mistakes are no problem. ・Any misalignment among top-1 to K (i.e., top) items. ・Any misalignment among top-K+1 to |A| (i.e., tail) items. ・Any misalignment whose probability ratio is bounded by the reward ratio: Alignment of the LSR’s action choice probability matters September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 85
  85. Theoretical analysis (3/3) 3) When CA-PG can learn the accurate

    alignment of actions? ・Any (oracle) epsilon-greedy and softmax-type policies satisfies the alignment condition of LSR. ・Even when the LSR policy makes mistake in the action alignment, the following mistakes are no problem. ・Any misalignment among top-1 to K (i.e., top) items. ・Any misalignment among top-K+1 to |A| (i.e., tail) items. ・Any misalignment whose probability ratio is bounded by the reward ratio: CA-PG works with a reasonably accurate (practical) LSR policy. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 86
  86. Key takeaways from theoretical analysis • Proposed method (CA-PG) is

    a partial PG of the vanilla PG. • CA-PG enables the credit-assignment within the candidate set, by considering the marginal prob of action is being selected in (one of) top-K. • CA-PG intentionally ignore from which candidate action come from, greatly reducing variance by modifying the action space from 𝑂(|𝐴|𝐾 ) to 𝑂(|𝐴|). • CA-PG can learn the accurate alignment with a practical choice of LSR. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 87
  87. A potential drawback of credit-assigned PG CA-PG requires some computational

    overhead to compute 𝜋(𝑆𝐾(𝑎)|𝑥). The gradient computation of CA-PG requires 𝑂(𝐾𝐿), while that of V-PG is 𝑂(𝐾). September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 88
  88. Our recommendation: TOP1-PG As a practical soluation, we suggest a

    simplified alternative called TOP1-PG, Use 𝑆1 instead of 𝑆𝐾 only for the gradient computation (i.e., we actually sample 𝐾 actions using ESR). September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 𝑂(𝐿) 89
  89. Our recommendation: TOP1-PG As a practical soluation, we suggest a

    simplified alternative called TOP1-PG, Use 𝑆1 instead of 𝑆𝐾 only for the gradient computation (i.e., we actually sample 𝐾 actions using ESR). 𝑂(𝐿) with the commonly use Plackett-Luce policy. (We will test the performance in experiments.) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 90
  90. Our recommendation: TOP1-PG As a practical soluation, we suggest a

    simplified alternative called TOP1-PG, Use 𝑆1 instead of 𝑆𝐾 only for the gradient computation (i.e., we actually sample 𝐾 actions using ESR). with the commonly use Plackett-Luce policy. 𝑂(𝐿) TOP1-PG is equivalent to CA-PG when • Using sampling-with-replacement (SwR) approximation to compute probability • Using a single model for selecting top-K actions (i.e., not using mixture-of-expert) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 91
  91. Pros and cons of PG methods Summarizing the properties of

    each PG, we have.. worst / best ※ SwR: Sampling-with-Replacement approximation MoE: mixture-of-experts September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 92
  92. Pros and cons of PG methods Summarizing the properties of

    each PG, we have.. worst / best In experiments, we test • How does the performance of each PG change with varying # of candidate (K), # of outputs (L), optimality of LSR? • How does the computational time change with varying # of candidates (K), # of outputs (L)? September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 93
  93. Additional results combined with GRPO [Shao+,24] CA-PG can be easily

    combined with other variance reduction methods. As a proof of concept, we combined CA-PG with GRPO: 1. Query 𝑚 samples per context and action, 𝑟𝑗(𝑥, 𝑎), 𝑗 ∈ [𝑚]. 2. Normalize the reward as 𝑟 ′ 𝑗 = (𝑟𝑗 − 𝑚𝑒𝑎𝑛(𝑟))/𝑠𝑡𝑑(𝑟). 3. (Add a constant value to scale rewards to be positive). September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 94
  94. Plackett-Luce and Sampling-with-Replacement (SwR) The PL policy selects candidate set

    by recursively applying softmax on the remaining. SwR approximation calculates the probability as if applying softmax independently. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 95
  95. Gradient Computation of CA-PG (1/2) The score function (log action

    choice probability) can be calculated as follows. We approximate this probability with small relative errors (~6%) in the next slides. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 96
  96. Gradient Computation of CA-PG (2/2) The probability is approximated as

    follows. Replacing the expectation over all possible candidate set with the most likely candidate set. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 97
  97. Appendix for the Orchestration paper An Off-Policy Learning Approach for

    Steering Sentence Generation towards Personalization. [KCSJ, RecSys2025] (PhD work at Cornell) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 98
  98. Theoretical analysis; support condition ① DSO is less likely to

    incur deficient support. (similar sentence support) because (action support) the similar sentence support is a relaxed condition of the action support. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 99
  99. Theoretical analysis; bias ② DSO has small bias when kernel

    bandwidth 𝛕 is small. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 100
  100. Theoretical analysis; bias ② DSO has small bias when kernel

    bandwidth 𝛕 is small. • This term comes from the within-neighbor reward shift. • • These terms comes from applying marginalization via kernels in the sentence space. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 101
  101. Theoretical analysis; variance ③ DSO has a large variance reduction

    when kernel bandwidth 𝛕 is large. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 102
  102. Theoretical analysis; variance ③ DSO has a large variance reduction

    when kernel bandwidth 𝛕 is large. • This term reduces variance by avoiding within-neighbor importance weights: • This term reduces variance by doing implicit data augmentation and soft-rejection sampling. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 103
  103. Synthetic experiments evaluation metric compared methods • Regression [Konda&Tsitsiklis,99] •

    IS [Swaminathan&Joachims,16] optimal policy uniform random • DR [Dudík+,11] the higher, the better • POTEC [Saito+,24] • DSO (ours) DR: hybrid of regression and IS POTEC: two-stage policy that uses the cluster of actions September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 104
  104. Synthetic experiments data generation process smooth, different prompts can results

    in similar sentence sentence • sentence distribution prompt reward • reward distribution smooth, different sentences results in different rewards September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN sentence 105
  105. Synthetic experiments configurations • data sizes: {500, 1000, 2000, 4000,

    8000} • number of candidate prompts: {10, 50, 100, 500, 1000} • reward noises: {0.0, 1.0, 2.0, 3.0} • For DSO, we use the Gaussian kernel with 𝜏 = 𝟏. 𝟎. value: default value September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 106
  106. Synthetic experiments ablations • kernel bandwidth: {0.5, 1.0, 2.0, 4.0}

    • logging marginal density: {w/ and w/o function approx.} (w/o is the monte-carlo estimation) • add noise 𝝈𝒔 = 𝟏. 𝟎 to the sentence embeddings to measure the distance value: default value September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 107
  107. Ablation results • We observe some bias-variance tradeoff when using

    monte-carlo estimation. • Using a Gaussian kernel and the function approx. improves the robustness of DSO to the choice of bandwidth hyperparameter 𝜏. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 108
  108. Why does function approx. improve the robustness of DSO? A.

    Because we use the MSE loss to fit the marginal density model. For example, when the true marginal density is 1e-5, estimating it as 1e-5 and 1e-4 does not change the MSE loss too much. In contrast, 1e-4 and 1e-5 make a significant difference. Using function approximation, we can avoid being too precise about small values of the marginal density. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 109
  109. Full-LLM experiment • Semi-synthetic experiments on the MovieLens-10M dataset [Harper&Konstan,15].

    • DistilBert [Sanh+,19]-based reward simulator is trained on the data. (next page) • User and query (i.e., movie) are sampled from the dataset. • Candidate prompts are retrieved from RelatedWord.io. • Using Mistral-7B [Jiang+,23] as the frozen LLM to generate the sentence. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 110
  110. Reward simulator fine-tuning on Movielens-10M Original CF dataset Augmented dataset

    • 𝑢: user id • 𝑞: item id (movie title) movie description • 𝑟: ratings (generated by Mistral-7B (zero-shot, w/o prompt)) Reward simulator user id embedding September 2026 inner product (・) DistilBert encoder movie description loss function: MSE in reward prediction End-to-End Training of Two-Stage Decision Systems @ UMN 111
  111. Examples of sentence generation in full-LLM bench. September 2026 End-to-End

    Training of Two-Stage Decision Systems @ UMN 112
  112. Reward simulation results of full-LLM bench. (Left) “positive” indicate the

    movies with a rating of 5, while “negative” indicates those with ratings of 0-3. (Right) Showing the distribution of normalized reward, which indicate the improvement of expected reward gained by using the given prompt, compared to that of the sentence generated without prompts. The normalized value is multiplied by 10, so the difference become evident when running policy learning methods. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 113
  113. How to leverage similarities among sentences? We consider estimating the

    following gradient in the sentence space. (true sentence policy gradient) gradient w.r.t. sentence distribution however, the issue is that the original sentence space is high dimensional.. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 114
  114. How to leverage similarities among sentences? We consider estimating the

    following gradient in the marginalized sentence space. (true marginalized sentence policy gradient) 𝜙(𝑠): kernel-based neighbors of sentence 𝑠 gradient w.r.t. marginalized sentence distribution September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 115
  115. How to leverage similarities among sentences? We consider estimating the

    following gradient in the marginalized sentence space. (true marginalized sentence policy gradient) 𝜙(𝑠): kernel-based neighbors of sentence 𝑠 gradient w.r.t. marginalized sentence distribution where September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 116
  116. How to leverage similarities among sentences? We consider estimating the

    following gradient in the marginalized sentence space. (true marginalized sentence policy gradient) 𝜙(𝑠): kernel-based neighbors of sentence 𝑠 gradient w.r.t. marginalized sentence distribution where (probability of observing sentence within 𝜙(𝑠)) (expected reward within 𝜙(𝑠) under policy π) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 117
  117. Direct Sentence Off-policy gradient (DSO) We estimate the sentence policy

    gradient using logged data as follows. (estimating marginalized sentence policy gradient) How we can actually estimate/implement the weighted score function? September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 118
  118. Estimation of the weighted score function We use the following

    re-sampling technique. ② ① ③ , which suggests that ① DSO does implicit data augmentation via resampling (𝑎, 𝑠′) from the policy 𝜋𝜃 . ② DSO uses soft rejection sampling using the kernel weight. ③ DSO corrects the logging distribution in the marginalized sentence space. See Appendix for the derivation. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 119
  119. How to estimate the logging marginal density? To use DSO,

    we need to estimate the logging marginal density defined as We can use function approximation trained on the following MSE loss. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 120
  120. Derivation of the weighted score function (1/2) As a preparation,

    we first derive the following expression of the importance weight. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 121
  121. Derivation of the weighted score function (2/2) Then, we transform

    the weighted score function as follows. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 122
  122. Naive approaches Naive approaches estimate the policy gradient (PG) to

    update the policy. regression-based [Konda&Tsitsiklis,99] important sampling-based [Swaminathan&Joachims,16] • impute regressed reward • correct the distribution shift to be unbiased • introduce bias when the regression is inaccurate • variance can be significantly high • (regression is often demanding due to partial reward and covariate shift) • (especially with a rich set of prompts) September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 123
  123. A baseline approach: Doubly Robust (DR) [Dudík+,11] DR uses regression

    as a control variate as follows, for the variance reduction purpose. control variate September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 124
  124. A baseline approach: POTEC [Saito+,24] POTEC considers two-stage policies to

    leverage clusters among prompts. 𝑐: cluster regression-based greedy (estimating cluster policy gradient) IS w.r.t. clustering space control variate Not leveraging the information about generated sentences. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 125
  125. Is it possible to define a DR-style variant of DSO?

    When defining a DR-style estimator, the baseline term should be as follows. However, estimating the gradient involves another importance sampling, and we cannot reduce variance by using this formulation. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 126
  126. Is it possible to define a DR-style variant of DSO?

    When defining a DR-style estimator, the baseline term should be as follows. However, estimating the gradient involves another importance sampling, It would be interesting to explore how we can efficiently combine the regression and DSO as a potential future work! and we cannot reduce variance by using this formulation. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 127
  127. References (1/2) [Ma et al., 2020] Jiaqi Ma, Zhe Zhao,

    Xinyang Yi, Ji Yang, Minmin Chen, Jiaxi Tang, Lichan Hong, Ed Chi. Off-policy Learning in Two-stage Recommender Systems. WWW, 2020. [Swaminathan&Joachims,16] Adith Swaminathan and Thorsten Joachims. Batch learning from logged bandit feedback through counterfactual risk minimization. JMLR, 2016. [Shao et al., 2022] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y.K. Li, Y. Wu, Daya Guo. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. 2024. [Dudík et al.,11] Miroslav Dudík, John Langford, and Lihong Li. Doubly robust policy evaluation and learning. ICML, 2011. [Saito et al.,24] Yuta Saito, Jihan Yao, and Thorsten Joachims. Potec: Off-policy learning for large action spaces via two-stage policy decomposition. 2024. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 129
  128. References (2/2) [Gao et al., 2022] Chongming Gao, Shijun Li,

    Wenqiang Lei, Jiawei Chen, Biao Li, Peng Jiang, Xiangnan He, Jiaxin Mao, Tat-Seng Chua. KuaiRec: A Fully-observed Dataset and Insights for Evaluating Recommender Systems. CIKM, 2022. [Harper&Konstan,15] F Maxwell Harper and Joseph A Konstan. The movielens datasets: History and context. TIIS, 2015. [Jiang wt al., 2023] Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, William El Sayed. Mistral 7b. 2023. [Sanh et al., 2019] Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. 2019. September 2026 End-to-End Training of Two-Stage Decision Systems @ UMN 130