Applied ML Engineer · Data Scientist
I enjoy the whole process of building AI systems: understanding the problem, working with data and models, getting it into use, and improving it after launch. That last part is where the interesting questions are, like whether it actually helps the people using it.
I recently finished my MS in Computer Science at Boston University. Right now I'm working on how reliably LLM judges agree with human evaluators, and how to measure whether an LLM system is useful to the people it serves.
Before that, five years at Singtel, Singapore's largest telecom, owning propensity models on a $2M BCG engagement and designing the holdout that measured their impact. At BU I led technical delivery of four client ML projects, including a retrieval system for the Boston Public Library, and co-authored a paper on deepfake detection and courtroom admissibility.
Right now, I am drawn to LLM and agent evaluation, retrieval and search, inference engineering, and red-teaming.
Boston University · GPA 3.67
University of California, San Diego · Provost Honors
Python · Anthropic and OpenAI judges · Cohen's κ · bootstrap CIs
Pre-registered replication on MT-Bench: across three judges and 100 pairwise comparisons, judge–human agreement ran 0.17 to 0.21 below the human–human baseline, and judges picked ties on 3 to 9% of pairs versus 23 to 25% for humans.
BGE-M3 dense and sparse in pgvector · Neo4j entity graph · RRF + metadata rerank · GraphRAG expansion · GPT-4o intent parsing and cited generation · Streamlit on Hugging Face Spaces
Led technical delivery of hybrid retrieval with cited answers over library collections: MRR 0.049 to 0.209 and Hit@10 0.029 to 0.429 on 35 held-out queries.
CycleGAN · WGAN-GP · conditional VAE · ResNet-18 · 41,051 radiographs
Tested three generators at five augmentation ratios under a 135:1 class imbalance; none improved KL4 recall.
PyTorch · CNN-Transformer · 164 Apple Watch recordings
Benchmarked five architectures with recording-level cross-validation; CNN-Transformer reached 90.0% ± 5.7% accuracy.
Zulal Akarsu, Maryan Rizinski, Ming Zhang, Kyung-Shick Choi, Deven Shah, Ittoop Shinu Shibu, Chan Woo Shin, Lou Chitkushev. International Conference on Intelligent Digital Forensics and Cybersecurity (IDFC 2026). Accepted.
MLH Best Use of Gemini API and infrastructure track, for ForTheCity, sidewalk accessibility validation
For CollabNet, research collaborator discovery over OpenAlex
Insurance Business
Customer Success
Python · PyTorch · scikit-learn · LightGBM · SQL · Spark · Databricks · Snowflake · MLflow · Airflow · Docker · RAG · pgvector · Neo4j · LangGraph