Data engineer with ~10 years of experience building large-scale data platforms on AWS and Azure — PySpark, Apache Iceberg lakehouses, and event-driven architectures. I contribute to open-source data tooling.
OpenLineage: the open standard for data lineage
- SSL context support for the Python HTTP transport: added an
ssl_contextoption toHttpConfigso clients can use custom CA bundles and mTLS (issue #3460, PR #4983 — under review).
SQLMesh: data transformation framework
- Fix: drop clustering key before dropping columns it references (issue #5813, PR #6095 — under review).
Dagster: data orchestration platform
- Fix execution hang when a skip is unblocked by an abandon:
plan_events_iteratornow runs the skip/abandon phases to a fixed point instead of a single pass, so a skip unblocked by an abandonment can no longer stall execution forever (issue #33661, PR #34238 — under review).
PyIceberg: Apache Iceberg's Python library
- Fix: stream record batches lazily in
to_arrow_batch_reader(): the batch reader materialized every file up front (~12.6x memory blowup); it now streams batches one at a time (issue #2407, PR #4028 — under review).
- Silent data loss in serverless Spark — arXiv paper (2026).