CV
I am a computational biologist and a data scientist working in predictive modeling, probabilistic programming, probabilistic forecasting, statistical modeling, and machine learning. I have a MSc in Bioinformatics and completing a second MSc in Data Science.
Education
MSc Statistical Data Analysis (Computational Track) Ghent University, Belgium · 2025 – ongoing Thesis: Multivariate Probabilistic Forecasting for Predicting Risk Events (in progress, expected defense early 2027) Benchmarking study comparing diffusion models against established models for multivariate probabilistic forecasting.
MSc Bioinformatics (Systems Biology Track) Ghent University, Belgium · 2021 – 2024 Thesis: Leveraging Probabilistic Programming in Causal Inference for Precision Medicine Framed causal inference within Pearl’s structural causal model framework and benchmarked do-calculus, propensity score matching, causal forests, and probabilistic programming (Julia, Turing.jl) on synthetic data before applying the method to a real-world NHANES case study on Vitamin D and depression.
BSc Molecular Biology and Genetics Bogazici University, Turkey · 2016 – 2021
Experience
Data Scientist at Gastromind (Remote) · May 2024 – December 2024
- Led a 3-person data quality team performing daily cleanup and validation across hundreds of Airtable records per day, including deduplication and schema-conformance checks for downstream SQL migration pipelines
- Performed day-to-day SQL querying across relational databases for reporting, data validation, and post-migration accuracy checks
- Used Python on an ad-hoc basis to reformat and reshape data to fit existing schemas during ingestion
- Participated in daily cross-functional standups spanning the general team, data science team, and database cleanup team
Bioinformatics Scientist Intern at BioLizard (Ghent) · Aug 2022 – Sep 2022
- Built a microbiome-based prognosis prediction pipeline for a client project in a confidential disease area
- Applied classical machine learning methods (scikit-learn, pandas, numpy) to process and model microbiome sequencing data, achieving classification accuracy in the 70-80% range
- Collaborated with client stakeholders as part of a consulting engagement
Skills
Programming Languages Python (proficient) · R (proficient) · Julia (Turing.jl, Optim.jl) · SQL (proficient) · Bash
Statistical and Probabilistic Methods Bayesian inference · Probabilistic programming · MCMC · Causal inference (do-calculus, propensity score matching, causal forests, heterogeneous treatment effects, R-learner, doubly robust estimation) · Multivariate probabilistic forecasting · Regression modelling · Epidemiology and clinical trial design
Machine Learning and Deep Learning Scikit-learn · XGBoost · LightGBM · PyTorch · TensorFlow · Supervised and unsupervised learning · SHAP · Hugging Face / LLMs · RAG
Bioinformatics Biopython · DESeq2 / Bioconductor · Statistical genomics · Familiarity with standard bioinformatics CLI tools (alignment, variant calling, omics pipelines)
Data and Infrastructure Pandas · NumPy · PostgreSQL · Airtable · ETL pipelines · Git · Docker · SSH and remote server management · HPC clusters · Linux/macOS CLI
Languages
English (full professional) · Turkish (native) · Dutch (A2, working towards B1)