Diogo Ribeiro (@DiogoRibeiro7)
Lead Data Scientist · Researcher · Invited Assistant Professor
Statistical modelling · Causal inference · Time series · Bayesian methods · Applied AI
I build data science systems and research software with an emphasis on transparent modelling, uncertainty, reproducibility, and decisions that can be defended from the evidence.
My work spans statistical machine learning, causal inference, forecasting, optimisation, experimental design, applied AI, and numerical computing. I prefer strong baselines and mathematically grounded models before adding unnecessary complexity.
Selected work
These are the repositories I would start with for a representative view of my technical work.
| Repository | Focus |
|---|---|
| causal-econometrics | Reproducible causal inference and econometrics in Python: experiments, quasi-experiments, causal ML, panel methods, sensitivity analysis, and marketing measurement. |
| gaussian-processes-r | Gaussian processes from first principles in R, including composable kernels, Cholesky-based exact regression, marginal-likelihood gradients, calibration, posterior sampling, and sparse extensions. |
| numerical-linear-algebra-r | Numerical linear algebra in R with explicit diagnostics for least squares, preconditioned conjugate gradients, randomized SVD, rank, residuals, and backward error. |
| geo-placebo-experiment | Placebo-based evaluation of geo-experiment estimators, false-positive rates, minimum detectable lift, synthetic control, DiD, conformal inference, and structural time series. |
| gb-rail-delay-dynamics | Event-based modelling of lateness accumulation, carried delay, and recovery within Great Britain rail journeys. |
| dml-sensitivity-experiment | Sensitivity analysis for double machine learning with benchmark-calibrated omitted-variable-bias bounds and known-truth experiments. |
Other public studies cover policy learning, synthetic control, interrupted time series and BSTS, heart-rate modelling, Iberian marginal emissions, text-watermark detection power, and reproducibility experiments in R.
Expertise
| Area | What I work on |
|---|---|
| Statistical and causal modelling | Regression, causal inference, experimental and quasi-experimental design, uncertainty, sensitivity analysis, and model evaluation. |
| Time series and forecasting | Temporal validation, structural time series, state-space models, forecasting, change detection, and operational decision support. |
| Bayesian methods | Gaussian processes, probabilistic modelling, posterior uncertainty, calibration, and simulation. |
| Applied AI | Retrieval-augmented generation, evaluation, serving, observability, and reliable deployment. |
| Numerical and research computing | Stable numerical methods, simulation, reproducible experiments, typed Python/R code, tests, and traceable outputs. |
| Software delivery | CI/CD, repository governance, documentation, reusable automation, and maintainable research software. |
Engineering and developer tooling
GitLab is also where I develop tooling around reproducibility and software delivery:
- gitlab-project-audit — project governance, CI/CD, repository hygiene, pipeline history, variables, tags, releases, and machine-readable audit reports.
- gitlab-ci-collection — reusable GitLab CI/CD components with explicit inputs, versioned interfaces, examples, and validation.
- gitlab-pipeline-profiler — explainable analysis of pipeline runtime, failures, retries, queueing, and critical-path jobs.
- dataleaf — a Hugo theme for technical, mathematical, and long-form publishing.
How I work
- Define the question first. Establish the estimand, failure modes, constraints, and decision rule before choosing an algorithm.
- Start with strong baselines. Use classical statistical and mathematical models where they answer the question well; add complexity when the evidence supports it.
- Treat uncertainty as part of the result. Calibration, missingness, leakage, drift, sensitivity, and operating thresholds belong in the analysis rather than after it.
- Make claims reproducible. Use tests, CI, explicit environments, provenance, documented assumptions, and machine-readable outputs.
- Build for review. Prefer focused components and changes that can be inspected, tested, and maintained.
Professional and academic work
My industry background is in lead-level data science and applied research, including health-tech, behavioural sensing, forecasting, statistical modelling, and production data systems.
I am also an Invited Assistant Professor (part-time) at the Faculty of Media Arts and Design (FMAD), Technical University of Porto.
Selected delivery outcomes include an 80% reduction in reporting costs, a 30% reduction in analytics processing time, and a €500K reduction in inventory value through forecasting and operational optimisation.
For a broader data science and applied AI portfolio, visit GitHub.
Contact
For professional, research, or technical collaboration: