organizations be successful with Data and AI • Mix of Data Engineers, Data Scientists, Machine Learning Engineers, Analytics Engineers & Analytics Translators • Based in Amsterdam & Eindhoven • Part of Xebia Data & AI • The best Payment Service Provider out there • Founded in 2004 by Adriaan Mol • Mission to simplify financial services by creating world-class products • Currently active for merchants in European Economic Area (EEA), Switzerland, and the United Kingdom
Scientist • Prone to human errors • Labor intensive Manual predictions are convenient Data Scientists ≠ Software Engineers • Data Scientist tend not to have traditional Software Engineering background • Tend to lack understanding of DevOps ML models are something else • Different than regular software artifacts • There is significant overlap Machine Learning models to production is hard
any limitations or strings attached. Each DS has their own Virtual Machine Each VM: • Has access to data • Is persistent • Can be configured to work with VSCode or PyCharm on your local machine • Can have the specs you need/want Step 1: Let’s explore the problem! Workbench
step, we’ll: • Use the Pipeline components that Google provides out of the box to setup a training pipeline. • Train our model and output Datasets and Models. • Deploy our model to an Endpoint so our model is available for consumption by downstream users. Step 2: Deploy it as if you’re Google Metadata Models Pipelines Datasets Endpoints
Automated Train/Test splits are nice, but also opaque and not very configurable ❌ Where is the model evaluation step? ❌ Everything disappears into one big “train the model” step, including preprocessing. ✅ We have an ML Pipeline ✅ It’s all code ✅ It can be scheduled and kicked off automatically ✅ Everything now is traceable
the Kubeflow API It integrates well with Vertex AI and can be easily customized Mollievert is our package with customized components to simplify and clarify our ML Pipelines Step 3: Now Mollie-fy it! Metadata Models Pipelines Datasets Endpoints
• Making it easier to go to production without lowering the bar • Empowering Data Scientists and ML Engineers with tooling • Defining a ‘Golden Path’ to production, but allowing customization if desired.