// About
Thales Pomari
MLOps Engineer · Data Engineering
// Bio
I started my career as a researcher: I'm the first author of a deep learning paper published at IEEE ICIP 2018 (International Conference on Image Processing) on image splicing detection. That scientific foundation shaped the rigorous, data-driven approach I carry to this day — later consolidated with a Postgraduate Specialization in Data Science.
Today I'm an MLOps Engineer at Housecall Pro (US, remote), building the infrastructure that takes ML models from development to production: end-to-end automated pipelines, CI/CD for reproducible model releases, monitoring and alerting. Before that, as a Tech Lead at a multinational insurance company, I led a squad of 5 to 8 data engineers over a ~900 TB data lake — solution architecture, CDC pipelines with Kafka, distributed Spark ETLs — and led the data engineering and MLOps work of the company's first AI project, plus the regulatory integration with Open Insurance (SUSEP, the Brazilian insurance authority).
My current stack: Apache Spark and Apache Kafka for distributed processing and streaming; AWS (Glue, S3, Lambda, DynamoDB, SageMaker) as my main platform, with solid experience on GCP (BigQuery, Dataproc, Dataflow); and productionization of open-source LLMs (Gemma, Llama, and Qwen), with model serving and prediction APIs.
What sets me apart is learning fast and picking the right technology for each problem's timeline and budget — no over-engineering. From my time as a Tech Lead, I carry the human side of engineering: clear communication across teams, unblocking colleagues, and making sure technical decisions are understood by everyone involved, not just engineers.
// Experience
MLOps Engineer
Housecall Pro
Infrastructure to deploy, monitor, and manage ML models in production: end-to-end automated pipelines (feature engineering, deployment, and continuous evaluation), CI/CD for reproducible model releases, and monitoring, logging, and alerting — partnering with data scientists and product teams.
Tech Lead - Data Engineering & MLOps
Digiage Technologies, assigned to a multinational insurance company (life insurance and investment products)
- Led a squad of 5 to 8 data engineers: technical and strategic team management, solution architecture, and project management over a ~900 TB data lake (AWS, medallion architecture, Glue Workflows, Jobs, and Data Catalog).
- Led the integration of a new policy sales and claims management system (life insurance with investment product): business rules review, data modeling and mapping, and ETL development into the lake standard.
- Maintained and evolved a CDC pipeline with Apache Kafka for customer, policy, and complaint data, processed and served through DynamoDB with APIs consumed by the customer service CRM.
- Deployed ETL observability from scratch: a Python library for Glue jobs with Lambda-triggered alerts in Microsoft Teams (error details and CloudWatch log link), moving the team from incident-driven detection to proactive resolution before users noticed issues.
- Led the data engineering and MLOps work of the company's first AI project, a decision-support solution for claims analysis: aggregating customer, policy, and complaint data, integrating with the classification model developed by the data science team, and delivering grounded recommendations in a Power BI dashboard.
- Managed the Open Insurance integration: receiving data from other insurers and submitting data in the regulatory standard required by SUSEP.
- Enabled the ML platform for data scientists: SageMaker AI Notebooks with lifecycle scripts integrated with the data lake.
- Created automated deployment from scratch with AWS CodeBuild: backups, artifact generation, automatic change documentation, and publishing to S3, replacing a fully manual process.
- Restructured the team's code governance, previously treated as simple backup: new repositories with CloudFormation templates (infrastructure as code), conventional commits, and a feature/release branching flow, replacing direct commits to dev.
MLOps Engineer (part-time)
Inovia
- Built the production pipeline for 3 open-source LLMs (Gemma, Llama, and Qwen) to productionize a document analysis solution.
- Developed prediction serving and management APIs, with model and version orchestration.
- Manage the cloud infrastructure serving models to external customers.
- Integrate generative AI into the engineering workflow (Claude and agents): code review, code quality and security, code standardization, and logging, reducing review time and documentation dependency. Implemented unit and end-to-end tests with Claude assistance.
Data Engineer
Lima Consulting
- Implemented and maintained CDPs (Adobe Experience Platform / Real-Time CDP) for large telecom, retail, and banking companies; the main project served a telecom carrier with roughly 80 million customers.
- Modeled XDM schemas with schema evolution and built batch data ingestion (BigQuery connection and CSV file mapping) and streaming ingestion, configuring client-side Kafka integrated with AEP.
- Used AEP APIs to query customer, campaign, and segment data.
- Developed analytical reports through Query Service (PostgreSQL with nested JSON structures): parsing data to track each customer's campaign touchpoints, reproducing Adobe's native graphical view with more precise filters.
- Created multichannel activation journeys in Adobe Journey Optimizer (email, push, and SMS), including domain IP warming strategies, and contributed to the native Zenvia-Adobe integration for SMS delivery, specifying endpoints and documenting the data exchange between systems.
- Built data pipelines with Airflow and Redshift across AWS and GCP environments.
Data Engineer - MLOps
Boa Vista SCPC, now Equifax Brazil
- Created an internal Feature Store for registering and managing models and variables (SQL and metadata, with edit and delete), with OAuth 2.0 authentication integrated with GCP and a Bootstrap frontend.
- Refactored similarity and segmentation variable calculations from BigQuery to Spark/Scala on Dataproc (10 to 15 node cluster): runtime reduced from days, with failures and restarts, to about 4 hours, cutting cloud costs.
- Developed REST APIs in Flask and Airflow DAGs with Python jobs for data processing and model and variable calculation.
- Deployed through CI/CD and maintained projects on GCP: Kubernetes, Dataflow, Bigtable, BigQuery, and Cloud Storage.
Data Engineer
Digiage Technologies
- Developed Java MapReduce routines, running on on-premise servers, to generate prepaid and postpaid customer indicators for the marketing team of a major telecom carrier, processing ETLs of roughly 4.5 billion rows, with per-environment algorithm tuning.
- Supported business rule definitions, mapping the paths between tables to reach the information requested by the marketing team.
- Performed loads and optimizations on Hive and Greenplum through Apache Sqoop, produced analyses and reports, and maintained SAS scripts.
// Stack
Languages
- Python
- Scala
- Java
- SQL
- Shell Script
Distributed Data & Streaming
- Apache Kafka
- Apache Spark
- Dataflow
- Hive
- MapReduce
Orchestration
- Apache Airflow
- AWS Glue Workflows
AWS
- Glue
- S3
- Lambda
- DynamoDB
- RDS
- SageMaker
- CloudWatch
- CodeBuild
- CloudFormation
GCP
- BigQuery
- Dataproc
- Dataflow
- Bigtable
- GKE/Kubernetes
- Cloud Storage
- Cloud Build
MLOps
- Model and LLM Productionization
- CI/CD
- Observability and Alerting
- Model Serving
- Prediction APIs
Applied AI
- Claude and Code Agents
- Gemma
- Llama
- Qwen
// Education
Postgraduate Specialization in Data Science
Instituto Federal de São Paulo (IFSP), Campinas
Technologist Degree in Systems Analysis and Development
Instituto Federal de São Paulo (IFSP), Campinas
// Certifications
Adobe Certified Expert - Adobe Real-Time CDP
Exams AD0-E600 and AD7-E601
// Publications
Image Splicing Detection Through Illumination Inconsistencies and Deep Learning
POMARI, T.; RUPPERT, G.; REZENDE, E.; ROCHA, A.; CARVALHO, T.
// Languages
Let's build something.
Available for data engineering and AI infrastructure projects.
Get in Touch →Campinas, SP, Brazil