Speaker
Abstract
Building machine learning pipelines to extract structured data from unstructured text is a popular problem within an unpopular development lifecycle. We’ll talk through how you can use LLMs so that your schema can `interrogate` structured data from your unstructured text data in a declarative and typesafe way.
The challenge of converting unstructured text data into structured, usable data is a well-known adversary to engineers, analysts, and data scientists alike. In the traditional paradigm, this task has been the exclusive domain of specialists, often requiring the creation of bespoke models for each data feature. Missed a feature? Let’s circle back next quarter.
In this talk we’ll see that Large Language Models are surprisingly effective at not only rote extraction of structured data from documents, but extracting derived information and doing so in a type safe way that adheres to your data model. We’ll show how Marvin’s AI Models, grounded in Pydantic, let you interrogate your data with your data model by combining the potent reasoning capabilities of AI with the type boundaries set by Pydantic. By letting developers build NLP pipelines solely with their data model’s schema, this lets engineers and analysts enjoy a declarative development experience with NLP.
We’ll go through real life applications of how LLMs are being used in production: structuring electronic health records data, developing custom entity extraction pipelines, generating synthetic data for test driven development, and automated schema normalization for data warehousing.
QCon New York 2023 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
ML in Practice Hosted by Sid Anand Chief Architect and Head of Engineering @DatazoomFrom the same track
Thursday 15 June
10:35 Dumbo / Navy Yard Session AI/ML PostgresML: Leveraging Postgres as a Vector Database for AI Montana Low Machine Learning w/ PostgresML With the growing importance of AI and machine learning in modern applications, data scientists and developers are constantly exploring new and efficient ways to store and analyze large amounts of data. 11:50 Dumbo / Navy Yard Session Search Needle in a 930M Member Haystack: People Search AI @LinkedIn Mathew Teoh Machine Learning @ LinkedIn LinkedIn's search functionality is one of its oldest capabilities, allowing members to search for people they know, or to discover new connections. 13:40 Dumbo / Navy Yard Session AI/ML Going Beyond the Case of Black Box AutoML Kiran Kate Senior Technical Staff Member @IBM Research Most AutoML tools are black-box tools. They offer no code/low code tools (UI/simple APIs) for practitioners to get started quickly. While this helps beginners, most experienced data scientists/ML practitioners often need more control. 14:55 Dumbo / Navy Yard Session ML in Practice Back to Basics: Scalable, Portable ML in Pure SQL Evan Miller Principal Statistics Engineer @Eppo (Creator of Evan's Awesome A/B Tools) Redshift has SageMaker. BigQuery begat BigML. Spark birthed Databricks. Every data warehouse is tightly coupled to a particular ML stack. 16:10 Dumbo / Navy Yard Session LLMs in the Real World: Structuring Text with Declarative NLP Adam Azzam AI Product Lead @Prefect Building machine learning pipelines to extract structured data from unstructured text is a popular problem within an unpopular development lifecycle.