data Statistical estimation and tensor decompositions Econometrics and tensors: Sales potential in large retail network HR and tensors: Most effective store for each salesman From tensors to Tensor Train Neural Networks (TTNN) Industrial AI and TTNN: key part between deep sensors and deep actuators
Traditional form: “object-property” tables (99% of todays’ research and applications). Multi-modality: “object - context of type A - context of type A - . . . - time - . . . - control of type U - control of type V - . . . ” Measurements - events counts or utility level (economic effect, max joy, total reward, and alike). 6W: Who does What for What by What Where and When. x - y - z - t - f(x, y, z, t). State 1 - State 2 - Year - Treaty type - Scale - Effect.
as tensors and statistical estimation Data: (very) sparse random measurements, organized in multi-way arrays (tensors) Estimation problem: Optimal values of statistical distribution parameters (e.g. β in censored zero-inflate Poisson). Common base for supervised, unsupervised, and reinforcement learning. Goal: Optimal statistical decisions based on comparison of alternatives, in time.
Oserved sales only for certain SKU, in subset of stores, and in particular days. Very sparse. Observations are subject of censoring by items availability (censored by control process). Statistical mixture of censored and non-censored Poisson counts. Result: Estimation of sales potential for all SKU, all stores, and time. Items redistribution and logistics algorithm. More tensor dimensions: Marketing actions, prices, motivation, competitors.
Different sales people performance only partially measured in subset of stores. People vary in their skills, education, team context, etc. Who is better where? Result: Estimation of personal salesman potential in various working conditions. Recommendations to HR, better planning (and hunting).
Data is sparse, random, noisy, and inclined by (our) control. Data cannot be used in decision making per se. Statistical Estimates are stable and complete (cover all contexts). Statistical models also cover unobserved cases.
Generalization of SVD (NMF) from 2D-matrices to multi-way tensors. Maximum likelihood with stochastic gradient descent (ADAM+) T.Kolda Math to start: T.G. Kolda, B. W. Bader. “Tensor Decompositions and Applications.” SIAM Review, 51(3), pp.455-500.
train Decomposition in the form of diagonal matrices. Matrices can be arbitrary: Tensor Train (Ivan Oseledets, Institute of Numerical Mathematics, Russia, 2011)
Train Neural Network Neural layers instead of matrix multiplications! Tensor Train Neural Network (TTNN, S.Terekhov, Neuroinformatics 2017, MIFI, Moscow) Between sensors and actuators: Learning of large amount of small neural networks (hyper-graph, K. Anokhin) instead of huge deep nets.
Key element is TTNN associative mediation (“mind”) between sensor systems and control actions. Formula: DLactuators = f(DLsensors + TTNNcore + ES/logic · Ψ) Author forecast: Exploding interest to expert systems and "traditional"logical/Bayesian AI for Ψ