Skip to content
View BNTechie's full-sized avatar
💭
Statistics everywhere!
💭
Statistics everywhere!

Block or report BNTechie

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
BNtechie/README.md

Senior Data Scientist

Data Scientist (8+ years) working at the intersection of statistical genetics, population genomics, and applied machine learning in biomedical research. MEng in Big Data Analytics, PhD in Computational Physics.

Population Genomics & Statistical Genetics

  • GWAS, polygenic risk scores (PRS), multi-ancestry analysis
  • PRS tools: SBayesR (via GCTB), LDpred2, PRS-CS, PLINK (--score)
  • Causal inference: Mendelian randomization-style approaches using pharmacogenomic variants (e.g. CYP2D6, CYP2C19, CES1) as genetic instruments
  • Martingale residual transformation for Cox-to-linear GWAS
  • Large-scale cohort analysis (e.g. iPSYCH2015, N=105,477; ~10,000 PGS computed from FinnGen/UK Biobank summary statistics)

Machine Learning & Predictive Modeling

  • Programming: Python, R, SQL, Bash
  • Statistical modeling: generalized linear models, multivariate regression, time-series analysis (scikit-learn, statsmodels, pandas, numpy)
  • Predictive modeling: LASSO, Random Forest, XGBoost, ensemble/Super Learner methods — applied to clinical prediction modeling in oncology trial data
  • Applied transformer models via the Hugging Face pipeline API (BERT, DistilBERT, Sentence-BERT, BERTopic) and PyTorch for GPU-accelerated inference on HPC systems — applying pretrained models, not training architectures from scratch

Infrastructure & Pipelines

  • HPC cluster pipelines, Snakemake (basic level)
  • Containerization: Docker, Singularity
  • Version control: Git
  • Data visualization: Matplotlib, Seaborn, ggplot2

Research & Project Experience

  • Phenotype QC, ancestry-stratified association testing, and PRS analysis across large genomic cohorts
  • Predictive modeling for clinical trial data, contributing to early-phase trial insight generation
  • Cross-institutional collaboration on pharmacogenomic causal inference projects

Soft Skills

  • Written and verbal communication in English
  • Collaboration in interdisciplinary, multicultural, and cross-institutional teams
  • Independent project management

Pinned Loading

  1. Predictive-modeling Predictive-modeling Public

    Credit card fraud detection, Breast cancer prediction, Wine quality prediction, Bank note authentication, prediction of attrition of employees, Stock prediction, etc

    Jupyter Notebook 1

  2. snakemake-ml-pipeline snakemake-ml-pipeline Public

    Python

  3. my-blog my-blog Public

    Jupyter Notebook

  4. LDSC_h2_documentation LDSC_h2_documentation Public

    Shell

  5. moodtrack-fullstack moodtrack-fullstack Public

    Jupyter Notebook