My research is about drawing trustworthy conclusions from large observational databases. Machine learning can estimate causal effects there, but it rarely says how uncertain those estimates are, and without valid confidence intervals, results can't support reliable conclusions. I'm working to close that gap in three ways: developing confidence intervals that remain valid when machine learning handles hundreds of thousands of features, extending them to account for unobserved confounding, and making the resulting estimators as precise as possible. The goal is methods that let social and health scientists make valid causal claims from large-scale register data.