Back to Basics: Scalable, Portable ML in Pure SQL

QCon New York 2023

Session ML in Practice

Back to Basics: Scalable, Portable ML in Pure SQL

Thursday Jun 15 / 02:55PM EDT, Dumbo / Navy Yard

Abstract

Redshift has SageMaker. BigQuery begat BigML. Spark birthed Databricks. Every data warehouse is tightly coupled to a particular ML stack. This is good for warehouse vendors – but leads to vendor lock-in, implementation complexity, and significant frictions when shuttling data to and from the ML engine.

When Eppo was trying to predict end-user behavior so that our clients could conclude their A/B experiments more quickly (the CUPED algorithm), we realized there was an opportunity to build something "so crazy it might just work" – a portable regression engine that did all of the heavy compute inside each warehouse using bog-standard ANSI SQL, and saved the hard matrix math for our client code. By leveraging each warehouse's ability to crunch columns quickly, concurrently, and *in situ*, we were able to perform complex ML estimation in far less time than it previously took just to egress that data to a dedicated ML engine.

In this talk, I will walk through the architecture of Eppo's portable, performant, privacy-preserving, multi-warehouse regression engine, and discuss the challenges with implementation as well as the quirks associated with each warehouse. My goal is to challenge prevailing industry assumptions about what's needed to build scalable ML systems and show a way forward where we start moving compute closer to the data – using standard tools that are right in front of us. Attendees can expect to leave with a solid understanding of how Eppo's system works and just enough knowledge to start building a similar system in their programming language of choice.

Topics

ML in Practice MLOps Data Architecture
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon New York 2023 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Thursday 15 June

10:35 Dumbo / Navy Yard Session AI/ML PostgresML: Leveraging Postgres as a Vector Database for AI Montana Low Machine Learning w/ PostgresML 11:50 Dumbo / Navy Yard Session Search Needle in a 930M Member Haystack: People Search AI @LinkedIn Mathew Teoh Machine Learning @ LinkedIn 13:40 Dumbo / Navy Yard Session AI/ML Going Beyond the Case of Black Box AutoML Kiran Kate Senior Technical Staff Member @IBM Research 14:55 Dumbo / Navy Yard Session ML in Practice Back to Basics: Scalable, Portable ML in Pure SQL Evan Miller Principal Statistics Engineer @Eppo (Creator of Evan's Awesome A/B Tools) 16:10 Dumbo / Navy Yard Session LLMs in the Real World: Structuring Text with Declarative NLP Adam Azzam AI Product Lead @Prefect