I'm a data engineer and M.S. Computer Science student at Northeastern University (graduating Spring 2027). I build reliable data models, transformation workflows, and backend data systems using SQL, Python, Snowflake, dbt, and PostgreSQL.
Most recently, I worked on AI and data platforms at SharkNinja, where I developed Snowflake/dbt models for consumer analytics and worked with enterprise and vendor datasets from Oracle, Palantir, Kyra, MyTake, and Meltwater. I also built a Snowflake Cortex thematic-analysis pipeline operating across 256M rows and caught a deduplication defect that was silently dropping records.
I'm especially interested in data engineering, analytics engineering, data quality, and AI-adjacent data platforms. I am currently seeking Fall 2026 co-op opportunities and Summer 2027 internships.
- Modeling analytics-ready datasets with dbt and dimensional data-modeling practices
- Building repeatable ETL/ELT workflows in Python and SQL
- Improving data quality through testing, validation, and idempotent processing
- Designing PostgreSQL-backed services and optimizing high-volume transactional systems
- Exploring reliable ways to use LLM operators inside production data pipelines
Snowflake · dbt · SQL · Python · Snowflake Cortex
- Built dbt models that standardized MyTake survey data for purchase intent, star ratings, and price sensitivity.
- Worked with enterprise and vendor data from Oracle, Palantir, Kyra, MyTake, and Meltwater to support consumer and market analysis.
- Added incremental loading, change-data capture, and soft-delete handling for repeatable refreshes.
- Developed a multi-pass thematic-analysis workflow with Snowflake Cortex operating across 256M rows of consumer feedback data, and fixed a deduplication path that silently dropped survey responses.
- Automated Salesforce content translation across 12 markets, reducing localization time by 60%, and built an OCR workflow that tagged 10,000+ images with 95% accuracy.
Graduate Research Assistant — Food ALERT
PostgreSQL · Python · FastAPI · pytest · GitHub Actions
- Designed a normalized PostgreSQL schema and stored procedures for a five-stage food-donation lifecycle.
- Built data-backed FastAPI services and deployed them through an automated CI/CD workflow.
- Developed a six-signal organization risk score that reduced manual review time by 60%.
- Maintained 80%+ test coverage across the API and data layers.
T-SQL · MySQL · SQLite · Data Synchronization
- Tuned stored procedures and indexes for a production POS system processing 1,000+ daily transactions across two UK retail chains.
- Reduced API latency by 35% through query and database optimization.
- Built offline caching and reconnect handling to preserve order data during intermittent connectivity.
I'm strengthening my production data-engineering portfolio through three focused builds:
- Batch ELT: Python → S3 → Snowflake → dbt, orchestrated with Airflow and tested in CI
- Streaming: Python → Kafka/Redpanda → Snowflake, with deduplication, schema-drift handling, and a dead-letter queue
- Distributed & ML data: PySpark feature pipelines with MLflow tracking, batch scoring, and drift monitoring
I'm also developing a research project on contract testing for non-deterministic LLM data transformations, with planned dbt and Airflow integrations.
Data Engineering & Warehousing
SQL · Python · Snowflake · dbt · ETL/ELT · Data Modeling · Incremental Models · Data Quality Testing
Databases
PostgreSQL · T-SQL · MySQL · MongoDB · SQLite
Enterprise & Vendor Data Ecosystems
Oracle · Palantir · Kyra · MyTake · Meltwater
Cloud, DevOps & Services
AWS (S3, EC2, Lambda) · Docker · GitHub Actions · CI/CD · FastAPI · REST APIs · Linux
AI-Assisted Data Systems
Snowflake Cortex · Claude API · scikit-learn · Sentence Transformers · Pinecone · OCR
Currently building with
Apache Airflow · Kafka/Redpanda · PySpark · MLflow
- ❄️ SnowPro Core Certified (2026)
- 📄 Co-author, Natural Disaster Management System, IEEE ICCCSMD 2024
- 💡 Co-filer, Patent Application No. 202441052914
- 🏆 Hackathon winner and 2nd-place finisher among 72 teams
I'm open to conversations about data engineering and analytics engineering opportunities, data-platform projects, and research on reliable LLM-powered pipelines.



