Posts
All the articles I've posted.
- 24 MIN READ•May 23, 2026
Single-Node Data Engineering: DuckDB, DataFusion, Polars, and LakeSail
Optimize single-node data engineering with DuckDB, DataFusion, Polars, and LakeSail. Compare architectures and learn when to transition to Dremio MPP.
DuckDBApache ArrowDataFusion - 20 MIN READ•May 23, 2026
An In-Depth Overview of the Apache Iceberg 1.11.0 Release
Apache Iceberg 1.11.0 delivers manifest list encryption, the new pluggable File Format API, credential lifecycle refreshes, and Spark/Flink improvements.
Apache IcebergData LakehouseOpen Table Format - 22 MIN READ•May 22, 2026
Open Table Format Benchmarks: Why They Require Critical Evaluation
An in-depth analysis of open table format benchmarks comparing Apache Iceberg, Delta Lake, and Apache Hudi, detailing the pitfalls of standard benchmarks and how to choose a format.
open table formatsapache icebergdelta lake - 27 MIN READ•May 22, 2026
Apache Iceberg SCD Type 2 and CDC Patterns: Building Historical Lakehouse Tables
A deep dive into implementing Slowly Changing Dimension Type 2 (SCD Type 2) patterns and Change Data Capture (CDC) pipelines on Apache Iceberg, using PySpark and Dremio.
apache icebergcdcscd type 2 - 25 MIN READ•May 22, 2026
Setting Up an AWS-Native Open Lakehouse: Querying Apache Iceberg with AWS Athena and AWS Glue Catalog
A comprehensive guide to building an open, high-performance lakehouse on AWS using Apache Iceberg, AWS Glue Catalog, Amazon S3, and S3 Tables, with query acceleration via the Dremio engine.
Apache IcebergAWS AthenaAWS Glue Catalog - 24 MIN READ•May 22, 2026
Apache Iceberg Catalogs Explained: REST, Glue, Hive Metastore, Polaris, Nessie, and Snowflake
A deep dive into Apache Iceberg catalog architecture, comparing REST catalogs, AWS Glue, Project Nessie, Polaris, and Snowflake. Learn catalog role, credential vending, and cross-engine configurations.
apache icebergcatalogsNessie - 24 MIN READ•May 22, 2026
Maintaining Apache Iceberg Tables: Compaction, Snapshot Expiration, and Orphan File Cleanup
An in-depth guide to orchestrating maintenance operations on Apache Iceberg tables, covering bin-packing, sort-based, Z-Order compaction, snapshot expiration, and orphan file removal, with query acceleration details for the Dremio engine.
Apache IcebergCompactionData Engineering - 24 MIN READ•May 22, 2026
Apache Iceberg with Spark: Create, MERGE, Upsert, and Evolve Tables End to End
A comprehensive developer guide to configuring Apache Spark with Apache Iceberg, executing transactional writes, and managing schema evolution.
apache sparkapache icebergdata engineering - 21 MIN READ•May 22, 2026
Building a Multicloud Agentic Lakehouse Reference Architecture
A reference architecture for building an open, multicloud Data Lakehouse optimized for AI Agents using Apache Polaris, Apache Iceberg, and Dremio.
data lakehouseagentic lakehouseapache polaris - 21 MIN READ•May 22, 2026
Common Misconceptions About Data Lakehouse and Apache Iceberg
Addressing common search queries and reader confusion about Data Lakehouse architectures, Apache Iceberg catalogs, partitions, and lock-in.
data lakehouseapache icebergdata engineering - 7 MIN READ•Apr 29, 2026
Migrating to Apache Iceberg: Strategies for Every Source System
Migrate to Iceberg from Hive, data warehouses, or raw files using in-place migration, full rewrite, or the zero-downtime view swap pattern.
migrating to Apache IcebergHive to Iceberg migrationIceberg migration strategy - 7 MIN READ•Apr 29, 2026
Hands-On with Apache Iceberg Using Dremio Cloud
A practical walkthrough of creating, querying, and optimizing Iceberg tables on Dremio Cloud, from account setup to AI-powered analytics.
Dremio Cloud Apache IcebergDremio Cloud tutorialIceberg hands-on