ͱ speed ʹඞཁͳཁૉΛཏతʹ࣮ ˞ਫ਼ࣗମଞϥΠϒϥϦͱಉఔ දݪจΑΓҾ༻ Table 1: Comparison of major tree boosting systems. System exact greedy approximate global approximate local out-of-core sparsity aware parallel XGBoost yes yes yes yes yes yes pGBRT no no yes no no yes Spark MLLib no yes no no partially yes H2O no yes no no partially yes scikit-learn yes no no no no no R GBM yes no no no partially no t choosing 216 examples per block balances the erty and parallelization. cks for Out-of-core Computation There are several existing works on parallelizing ing [22, 19]. Most of these algorithms fall in proximate framework described in this paper. is also possible to partition data by columns [2 tree ͷׂʹ ࡍͯ͠Մೳͳ ߹ͤΛશ୳ࡧ ಛྔΛࢄԽͯۙ͠ࣅతʹѻ͏ global શͯͷࢬͰಉ͡ѻ͍ local split ຖʹѻ͍Λมߋ ϝϞϦʹΒͳ͍߹ ʹ֎෦ഔମ͔Βͷ ಡΈࠐΈͰಈ࡞ ࢄॲཧͷ࣮ εύʔεͳมʹର͢Δ efficient ͳॲཧͷ࣮ yes 2/10
ͷج४తมͷ࠷খԽ ɹ ※తؔ൚ؔͰ͋Γ Taylor ల։ͷೋ࣍ۙࣅΛ༻ → ֤มͷ֤ͰޯΛܭࢉ → શͯͷՄೳͳ split Λޮྑ͘୳ࡧ͢ΔͨΊʹมͷͰ sort ਤݪจΑΓҾ༻ Block structure for parallel learning. Each column in a block is sorted by the correspond linear scan over one column in the block is su cient to enumerate all the split points. 3/10
ʢม͕গͳ͍ঢ়گͰͷʣ௨ৗͷྨੑೳݕূ ɾYahoo LTRC : ϥϯΫֶशͷੑೳݕূ ɾCriteo : σʔλαΠζ͕େ͖͍ͷͰࢄॲཧͷੑೳݕূ දݪจΑΓҾ༻ Table 2: Dataset used in the Experiments. Dataset n m Task Allstate 10 M 4227 Insurance claim classification Higgs Boson 10 M 28 Event classification Yahoo LTRC 473K 700 Learning to Rank Criteo 1.7 B 67 Click through rate prediction We used four datasets in our experiments. A summary of these datasets is given in Table 2. In some of the experi- ments, we use a randomly selected subset of the data either due to slow baselines or to demonstrate the performance of the algorithm with varying dataset size. We use a su x to denote the size in these cases. For example Allstate-10K means a subset of the Allstate dataset with 10K instances. The first dataset we use is the Allstate insurance claim dataset8. The task is to predict the likelihood and cost of Table 500 t Meth XGB XGB sciki R.gb Ϩίʔυ มͷ 4/10
→ ޯͷऔಘͰඇ࿈ଓతͳϝϞϦΞΫηε͕ൃੜ͋͋͋ → େྔσʔλͰͷ֬อͷͨΊʹ cache-aware access Λ࣮ ਤݪจΑΓҾ༻ 1 2 4 8 16 Number of Threads 8 16 32 64 128 Time per Tree(sec) Basic algorithm Cache-aware algorithm (a) Allstate 10M 1 2 4 8 16 Number of Threads 8 16 32 64 128 256 Time per Tree(sec) Basic algorithm Cache-aware algorithm (b) Higgs 10M 1 2 4 8 16 Number of Threads 0.25 0.5 1 2 4 8 Time per Tree(sec) Basic algorithm Cache-aware algorithm (c) Allstate 1M 1 2 4 8 16 Number of Threads 0.25 0.5 1 2 4 8 Time per Tree(sec) Basic algorithm Cache-aware algorithm (d) Higgs 1M igure 7: Impact of cache-aware prefetching in exact greedy algorithm. We find that the cache-miss e↵ec mpacts the performance on the large datasets (10 million instances). Using cache aware prefetching improve he performance by factor of two when the dataset is large. 6/10 σʔλྔ͕ଟ͘ͳ͚Εখ͞ͳࠩ σʔλྔ͕ଟ͚Εݦஶͳࠩ
split finding Λ࣮ࢪ → missing Ͱͳ͍ཁૉʹରͯ͠ઢܗ࣌ؒͰಈ࡞ ߹Θͤͯ split ͷࡍʹ͕ͳ͍߹ͷ default ܾఆ ਤݪจΑΓҾ༻ Dataset: Allstate h column in a block is sorted by the corresponding feature is su cient to enumerate all the split points. 1 2 4 8 16 Number of Threads 0.03125 0.0625 0.125 0.25 0.5 1 2 4 8 16 32 Time per Tree(sec) Sparsity aware algorithm Basic algorithm Figure 5: Impact of the sparsity aware algorithm 8/10 50ഒఔͷ վળ