Causal effects, not just predictions
Estimating heterogeneous treatment effects (CATE/HTE) with doubly-robust and orthogonal methods, and using them to build decision policies rather than stop at a point prediction.
I work at the boundary between causal inference and machine learning: estimating the effect of actions, learning policies from those estimates, and building systems that decide under uncertainty. In production, that means contextual bandits and treatment-effect models inside large-scale personalization and marketing platforms. Alongside that, I'm building an independent research track, one paper at a time.
I started as a software engineer, moved into data science, and have spent the last few years narrowing in on the layer where statistics turns into action: causal inference, contextual bandits, and policy learning. Most of my day-to-day is building and operating decision systems in production — pricing, personalization, marketing — but I treat that work as a source of research questions, not just delivery. Competitions, an undergraduate thesis in space-plasma physics, and problems from work have each, in turn, become a methods question worth writing up properly.
A tabular Q-learning agent training live in your browser — no server, no dataset, just the Bellman update running client-side.
Episode 0 · ε = 1.00 · best path: —
My focus is the gap between a good model and a good decision: how to move from an estimate to a policy that can be trusted, measured, and defended once it's live.
Estimating heterogeneous treatment effects (CATE/HTE) with doubly-robust and orthogonal methods, and using them to build decision policies rather than stop at a point prediction.
Contextual bandits and Thompson Sampling for sequential decisions, with an eye on what breaks off-policy evaluation once a system starts learning from its own exploration.
A gradient-boosted regressor for realized takeoff-to-landing flight time that treats convective weather as geometry — segmenting storm imagery into georeferenced polygons and spatially joining them against radar-derived flight tracks, instead of collapsing weather into airport-level scalars. Grew out of a 2nd-place finish in the LATAM/ITA/DECEA air-traffic data science challenge.
A reusable library inspired by ESA Cluster mission data. Originated from an undergraduate thesis on magnetic reconnection and extreme-event statistics, now released as a tested package with a clear API.
Production-oriented implementation and backtesting utilities for portfolio optimization with Brazilian equities.
Data Mining Cup solution pipeline; documented approach and experiments.
Documented solutions with complexity notes and heuristic strategies.
Selected repositories: morepizza_hashcode, hashcodeDataCenterOptimization, adventofcode20.
Shows how combining SVD-based user and category similarity with Markov chains forecast weekly repeat purchases for DataMiningCup 2022, using Bayesian tuning to hit 92% accuracy within hardware limits.
A practical walkthrough on detecting heavy-tail anomalies; statistical assumptions, implementation notes, operational caveats, and visualization tips.
Happy to talk about causal inference, bandits, decision systems, or possible collaborations — research or applied. Reach out via the form or connect on LinkedIn.