Posts
All the articles I've posted.
- 13 MIN READ•May 24, 2026
Designing Governed RAG on Data Products
Enterprise RAG architecture that trusts its own data requires governance at the retrieval layer. Learn how to build governed RAG using data products, access policies, and semantic layer routing.
Governed Rag Enterprise Data ProductsEnterprise Rag ArchitectureGoverned Retrieval Augmented Generation - 12 MIN READ•May 24, 2026
What Iceberg V3 Advances Mean for CDC Pipelines
Apache Iceberg V3 brings deletion vectors and row lineage that reshape CDC pipeline design. Learn what these features mean for your streaming data architecture.
Iceberg Cdc PipelineIceberg Deletion VectorsIceberg Row Lineage - 13 MIN READ•May 24, 2026
Kafka 4.0 Changes Streaming Platform Operations
Kafka 4.0 removes ZooKeeper and ships KRaft and KIP-848 by default. Learn what those changes mean for platform operations, upgrades, and client configurations.
Kafka 4.0 UpgradeKafka KraftZookeeper Removal Kafka - 13 MIN READ•May 24, 2026
Lance and Iceberg for Multimodal AI Data
LanceDB and Apache Iceberg serve complementary roles in a multimodal AI lakehouse. Learn when to use Lance for embeddings and random access, and Iceberg for structured metadata and SQL analytics.
Lancedb Iceberg Multimodal Ai DataLancedb FormatLance Vs Iceberg - 13 MIN READ•May 24, 2026
Bringing MLflow and Data Pipelines Closer Together
MLflow 3 extends observability from classic ML experiments to GenAI tracing and data pipeline lineage. Learn how to connect data quality monitoring with model performance tracking.
Mlflow Data Pipeline ObservabilityMlflow 3 Genai TracingMlflow Data Quality Monitoring - 14 MIN READ•May 24, 2026
Modern Feature Stores Beyond Batch Pipelines
Feature stores like Feast now support streaming feature views from Kafka and Kinesis alongside batch pipelines. Learn how to build real-time features that maintain training-serving consistency.
Feature Store Streaming Real-Time MlFeast Streaming FeaturesOnline Offline Feature Store - 12 MIN READ•May 24, 2026
OpenLineage as the Spine of Data Observability
OpenLineage provides a standard API for collecting pipeline lineage across Airflow, Spark, Flink, and dbt. Learn how it powers blast radius analysis and incident triage.
Openlineage Data ObservabilityOpenlineage AirflowOpenlineage Spark - 12 MIN READ•May 24, 2026
When Paimon Beats Iceberg for Mutable Streams
Apache Paimon uses LSM-Tree storage for native CDC upserts without restart. Learn when Paimon outperforms Iceberg for high-churn mutable streaming workloads.
Apache Paimon Mutable StreamsPaimon Vs IcebergCdc Streaming Lakehouse - 13 MIN READ•May 24, 2026
Policy as Code for Lakehouse Governance
OPA, ABAC, row filters, and column masks make lakehouse governance programmable and scalable. Learn how Databricks, Snowflake Horizon, and BigQuery implement policy-as-code.
Policy As Code Data Governance LakehouseOpa Rego LakehouseAbac Databricks - 11 MIN READ•May 24, 2026
Real-Time Lakehouse Patterns with Apache Flink and Iceberg
Learn how to build a real-time lakehouse with Apache Flink 2.1 and the Dynamic Iceberg Sink, covering schema evolution, exactly-once delivery, and compaction.
Real-Time Lakehouse FlinkFlink Iceberg SinkKafka To Iceberg - 14 MIN READ•May 24, 2026
Why Semantic Layers Make Enterprise Text-to-SQL Safer
Text-to-SQL accuracy jumps from 40% to 85-95% when grounded in a semantic layer. Learn how Dremio, Snowflake Cortex Analyst, and dbt Semantic Layer improve AI analytics reliability.
Semantic Layer Text-To-SqlDremio Semantic LayerSnowflake Cortex Analyst - 14 MIN READ•May 24, 2026
Choosing Vector Stores for Retrieval Workloads
pgvector, Milvus, Weaviate, and LanceDB each make different tradeoffs on index type, hybrid search, scale, and operational complexity. Learn which fits your retrieval workload.
Vector Store Comparison Retrieval WorkloadsPgvector HnswMilvus Hybrid Search