Digiage Technologies · multinational insurance company · 06/2024 - 08/2026
Data Lake and ETL Observability at Scale
// Problem
A constantly growing ~900 TB data lake, with multiple squads producing ETLs, had purely reactive failure detection — issues only surfaced when a business user reported incorrect or missing data, often days after the original failure.
// Solution
- → Led a squad of 5 to 8 engineers on lake architecture and governance, standardizing medallion layers and Glue Workflows/Jobs.
- → Built observability from scratch: a shared Python library for Glue jobs, with Lambda-triggered alerts in Microsoft Teams (error details + direct CloudWatch link).
- → Standardized triage: every failure reaches the team with enough context for immediate diagnosis.
// Impact
Reactive incident-driven detection → proactive resolution, ahead of end-user perception.