Choosing Stream Pipelines
Stream Pipelines provides the most value when the source application publishes data through Streaming Ingestion. Streaming Ingestion pushes events to the pipeline immediately when events occur in the source system.
Stream Pipelines can process data that batch ingestion methods publish, including the Batch Ingestion API and ION. Batch ingestion methods store data in Data Lake before Stream Pipelines processes the data. Data Lake storage adds latency to data delivery.
Customers who rely primarily on batch ingestion methods must not expect real-time data delivery.
Consider the following use cases when you decide whether Stream Pipelines is the appropriate tool.
Use Stream Pipelines in the following scenarios:
- The source application uses Streaming Ingestion, and near-real-time data delivery to a cloud database or data warehouse is required.
- Operational reporting or event-driven processes require data in an external system with minimal delay.
Use an alternative method in the following scenarios:
- The requirement includes large-scale batch extraction or periodic loading of historical data. For these scenarios, use the Compass query platform, Compass APIs, or the Data Lake Objects API to extract data from the Data Lake and load data into the target system.
- The source application does not use Streaming Ingestion, and higher latency is acceptable. The Data Fabric ETL Tool remains suitable for batch-oriented extraction workloads.
Stream Pipelines do not support delivery to on-premises databases. Stream Pipelines do not support private connectivity methods, including AWS PrivateLink and Azure Private Link. Destination systems must be accessible through the internet by using a JDBC connection. Snowflake destinations must use the Snowflake Streaming endpoint.