Streaming from Apache Iceberg - Building Low-Latency and Cost-Effective Data Pipelines

QCon New York 2023

Session Stream Processing

Streaming from Apache Iceberg - Building Low-Latency and Cost-Effective Data Pipelines

Tuesday Jun 13 / 11:50AM EDT, Salon E

Abstract

Apache Flink is a very popular stream processing engine featuring sophisticated state management, even-time semantics, exactly-once state consistency. For low latency processing, Flink jobs typically consume data from streaming sources like Apache Kafka. Apache Iceberg is a widely adopted data lake technology supporting numerous features like snapshot isolation, transactional commit, fast scan planning. While Iceberg was originally designed for batch, it can also be used as a streaming source in Flink. This not only lowers the processing delays from hours or days to just minutes, but also significantly reduces the infrastructure cost and operational burden.

In this talk, we will explain the design of the Flink Iceberg source that we contributed to Apache Iceberg open source project. We will compare the Kafka and Iceberg sources for streaming read and present performance evaluation results of the Iceberg streaming read. We will discuss how the Iceberg streaming source can power many common stream processing use cases (like ETL, feature engineering). It enables users to build low-latency streaming pipelines chained by Iceberg that are cost effective and easy to operate.

Topics

Stream Processing Apache Flink Apache Iceberg Data Pipelines Architecture
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon New York 2023 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 13 June

10:35 Salon E Session Streaming Laying the Foundations for a Kappa Architecture - The Yellow Brick Road Sherin Thomas Staff Software Engineer @Chime 11:50 Salon E Session Stream Processing Streaming from Apache Iceberg - Building Low-Latency and Cost-Effective Data Pipelines Steven Wu Software Engineer @Apple and Apache Iceberg PMC 13:40 Salon D Session Serverless The Rise of the Serverless Data Architectures Gwen Shapira Founder @Nile, PMC Member @Kafka 14:55 Salon E Session Data Architecture Building a Large Scale Real-Time Ad Events Processing System Chao Chu Software Engineer @DoorDash 16:10 Salon E Session Architecture Enabling Remote Query Execution Through DuckDB Extensions Stephanie Wang Founding Engineer @MotherDuck 17:25 Carroll Gardens Unconference Unconference: Modern Data Architecture & Engineering Ben Linders Independent Consultant in Agile, Lean, Quality and Continuous Improvement