requires… collecting enough data careful user assignments to avoid overlapping experiments dealing with uncertainty Experiment cancellation, re-iteration, rescheduling releasing resources for other experiments starting experiments as soon as possible for every experiment Variants of the application !3
assignments to avoid overlapping experiments dealing with uncertainty Experiment cancellation, re-iteration, rescheduling releasing resources for other experiments starting experiments as soon as possible for every experiment 97% 3% v1 v2 Monitoring data Variants of the application Optimization Problem Goal: Identifying a valid schedule for executing multiple experiments with maximal fitness !4
"minDuration": 168, “sampleSize": 1000000, "priority": 8, "preferredUserGroup": [ "group3", "group5" ] } regression or business constant or gradual unique interactions Schedule: Schedule Traffic Assignment for Hour 2 of Experiment 4 Execution Plan for Experiment 4 Start Slot ! A 1 A 2 A 3 A N 24 4 % 2 % 0 % UG 1 UG 2 UG 3 0 % UG N … … Exec.Plan Exp 1 Exec.Plan Exp 2 Exec.Plan Exp 3 Exec.Plan Exp 4 Exec.Plan Exp N-1 Exec.Plan Exp N … !7
"minDuration": 168, “sampleSize": 1000000, "priority": 8, "preferredUserGroup": [ "group3", "group5" ] } regression or business constant or gradual unique interactions Schedule: Schedule Traffic Assignment for Hour 2 of Experiment 4 Execution Plan for Experiment 4 Start Slot ! A 1 A 2 A 3 A N 24 4 % 2 % 0 % UG 1 UG 2 UG 3 0 % UG N … … Exec.Plan Exp 1 Exec.Plan Exp 2 Exec.Plan Exp 3 Exec.Plan Exp 4 Exec.Plan Exp N-1 Exec.Plan Exp N … Constraints: same user groups during all time slots [business experiments only] sufficient data points for every time slot non-interrupted experiments sufficient traffic available for every time slot 1 2 3 4 !8
"minDuration": 168, “sampleSize": 1000000, "priority": 8, "preferredUserGroup": [ "group3", "group5" ] } regression or business constant or gradual unique interactions Schedule: dsi = minDuration scheduledDuration ssi = 1 τ usi = ∑d 1 coverage(Ai ) scheduledDuration f = wss * ∑n 1 ssi * pi ∑n 1 pi + wds * ∑n 1 dsi * pi ∑n 1 pi + wus * n ∑ 1 usi * pi Schedule Traffic Assignment for Hour 2 of Experiment 4 Execution Plan for Experiment 4 Start Slot ! A 1 A 2 A 3 A N 24 4 % 2 % 0 % UG 1 UG 2 UG 3 0 % UG N … … Exec.Plan Exp 1 Exec.Plan Exp 2 Exec.Plan Exp 3 Exec.Plan Exp 4 Exec.Plan Exp N-1 Exec.Plan Exp N … Constraints: same user groups during all time slots [business experiments only] sufficient data points for every time slot non-interrupted experiments sufficient traffic available for every time slot 1 2 3 4 1 2 3 Start score: Duration score: User group score: Combined Fitness Score: weighting Start Duration User Group Fitness [weighted-sum strategy]: !9
[business experiments only] sufficient data points for every time slot non-interrupted experiments sufficient traffic available for every time slot 1 2 3 4 Fitness [weighted-sum strategy]: priority of experiment i Experiment: { "id": 4, "type": "REGRESSION", "baseType": "GradualExperiment", "minDuration": 168, “sampleSize": 1000000, "priority": 8, "preferredUserGroup": [ "group3", "group5" ] } regression or business constant or gradual unique interactions Schedule: Schedule Traffic Assignment for Hour 2 of Experiment 4 Execution Plan for Experiment 4 Start Slot ! A 1 A 2 A 3 A N 24 4 % 2 % 0 % UG 1 UG 2 UG 3 0 % UG N … … Exec.Plan Exp 1 Exec.Plan Exp 2 Exec.Plan Exp 3 Exec.Plan Exp 4 Exec.Plan Exp N-1 Exec.Plan Exp N … dsi = minDuration scheduledDuration ssi = 1 τ usi = ∑d 1 coverage(Ai ) scheduledDuration f = wss * ∑n 1 ssi * pi ∑n 1 pi + wds * ∑n 1 dsi * pi ∑n 1 pi + wus * n ∑ 1 usi * pi 1 2 3 Start score: Duration score: User group score: Combined Fitness Score: Start Duration User Group !10
2 3 4 Genetic Algorithm: Mimics evolutionary process Reproduction steps within each generation: 1 2 3 4 Parent selection [Fitness Proportionate Selection] Crossover Offspring mutation Fitness and validity evaluation Next generation selection 5 Initial population created using random sampling Chromosome representation: Chromosome Traffic Assignment for Hour 2 of Experiment 4 Execution Plan for Experiment 4 Start Slot ! A 1 A 2 A 3 A N 24 4 % 2 % 0 % UG 1 UG 2 UG 3 0 % UG N … … Exec.Plan Exp 1 Exec.Plan Exp 2 Exec.Plan Exp 3 Exec.Plan Exp 4 Exec.Plan Exp N-1 Exec.Plan Exp N … !11
Annealing 1 2 3 4 Find valid solutions by creating individuals by chance Fitness function for assessing the individuals Constraints for checking individual’s validity Created starting population for GA, local search, and simulated annealing Local Search: Pick best individual of starting population Apply mutation operations of genetic algorithm If resulting fitness score higher, then replace current solution Repeat for multiple iterations Simulated Annealing: Similar to local search Take solutions with worse fitness with a certain probability Escape local optima !13
Real-World traffic profile [GitLab*], traffic divided into 5 user groups * https://monitor.gitlab.net/dashboard/db/fleet-overview Variations for required experiment sample sizes (RESS): low [15 million data points] medium [30 million data points] high [55 million data points] 10 baseline experiments [1 to 18 days] 6 regression-driven [2 gradual, 4 constant traffic] 4 business-driven 1 2 3 Duplicated 10 baseline experiments to create sets of 15, 20, 25, …, 70 experiments in 3 variants (low, medium, high RESS) each Parameter calibration GA population size, number of generations, crossover/mutation probabilities, LS/SA number of iterations, … Evaluation on Google Compute Engine 0 100,000 200,000 300,000 400,000 500,000 0 500 1000 Schedule Duration [hours] Traffic [unique requests] Total traffic low RESS medium RESS high RESS Traffic profile for user group 3 and traffic consumption of 3 example schedules (30 experiments each) with low, medium, and high RESS !15
schedule for 30 experiments after 72 hours) Reevaluation takes into account finished, canceled, newly added experiments Reevaluation adjusts (required) sample sizes of running experiments => actual traffic profile vs. estimation => adapt schedules based on gained knowledge Evaluation setup: Select best schedule of GA (74%) for 30 experiments with medium RESS Conduct reevaluation after 72 hours: 3 canceled experiments 3 finished experiments within 72 hours 5 new experiments with medium RESS added 1 2 10 runs in total for every approach • • • • 0.55 0.60 0.65 0.70 0.75 Random Sampling Genetic Local Search Simulated Annealing Fitness Reevaluation Score 30 Exp. Average Genetic Algorithm 73% 68% Local Search 69% 50% Simulated Annealing 67% 51% 1 2 Smaller gap between fitness scores of approaches [mean fitness] Execution times on similar level !18
even on simple cloud instances Further parallelization possible Importance of calibration Fine-tuning of parameters for different numbers of experiments Weighting of objectives (start score, duration score, user group score) Crossover Current “greedy” crossover leads to many invalid schedules Space for improvement taking validity constraints better into account !19