Ittoop Shinu Shibu
Ittoop Shinu Shibu

Ittoop Shinu Shibu

Applied ML Engineer · Data Scientist

About

I enjoy the whole process of building AI systems: understanding the problem, working with data and models, getting it into use, and improving it after launch. That last part is where the interesting questions are, like whether it actually helps the people using it.

I recently finished my MS in Computer Science at Boston University. Right now I'm working on how reliably LLM judges agree with human evaluators, and how to measure whether an LLM system is useful to the people it serves.

Before that, five years at Singtel, Singapore's largest telecom, owning propensity models on a $2M BCG engagement and designing the holdout that measured their impact. At BU I led technical delivery of four client ML projects, including a retrieval system for the Boston Public Library, and co-authored a paper on deepfake detection and courtroom admissibility.

Right now, I am drawn to LLM and agent evaluation, retrieval and search, inference engineering, and red-teaming.

Experience

Aug 2022 – Aug 2025

Data Scientist

Singtel, Singapore XGBoost · LightGBM · Llama 4 RAG · Databricks · Spark · MLflow · Snowflake clean room
  • Built and deployed propensity and customer-product alignment models on Singtel and Etiqa insurance data in a Snowflake clean room during a $2M BCG engagement: +80% YoY lead conversion, +210% product-line revenue growth.
  • Designed the rollout measurement: A/B tests against a random 10% no-touch holdout held for three months; lifted cross-sell conversion +10% MoM.
  • Piloted a Llama 4 RAG system for analyst reporting: 86% agreement with analysts, 50% less reporting time.
  • Ran models as Databricks jobs (Spark, MLflow) with scheduled retraining, drift monitoring, and pre-release holdout validation.
Aug 2020 – Aug 2022

Business Analyst, promoted to Manager

Singtel, Singapore Survival models · regularized regression · Python · R · Airflow · Hadoop
  • Built a customer lifetime value model in Python and R that improved segment profitability by 15%.
  • Maintained Airflow pipelines on a Hadoop cluster, cutting delivery timelines ~20%.
  • Managed four analysts, raising team productivity by 30%.
Jun – Aug 2026

Special Initiatives Intern

BU Spark!, Boston University
  • Scoped 6+ client ML projects; prototyped an AI Chinese-character learning app and an IATI data pipeline (Python ETL, Neo4j).
Mar – Jun 2026

Graduate Research Assistant, Deepfake Forensics

Boston University
  • Built frame-extraction and dataset-standardization pipelines for DF40, FaceForensics++, Celeb-DF v2, and DeepSpeak v2, and evaluated detection models on A40 and H100 GPUs.
  • Co-author on a deepfake-forensics paper accepted at IDFC 2026; see Publications.
Jan – May 2026

ML Technical Project Manager

BU Spark!, Boston University
  • Led technical delivery of four client ML projects with teams of 4 to 6, including the Boston Public Library retrieval system below.

Education

Aug 2026

MS, Computer Science

Boston University · GPA 3.67

2020

BS, Data Science

University of California, San Diego · Provost Honors

Selected Work

2026

LLM-as-Judge Reliability Audit

Python · Anthropic and OpenAI judges · Cohen's κ · bootstrap CIs

Pre-registered replication on MT-Bench: across three judges and 100 pairwise comparisons, judge–human agreement ran 0.17 to 0.21 below the human–human baseline, and judges picked ties on 3 to 9% of pairs versus 23 to 25% for humans.

2026

Hybrid Retrieval for the Boston Public Library

BGE-M3 dense and sparse in pgvector · Neo4j entity graph · RRF + metadata rerank · GraphRAG expansion · GPT-4o intent parsing and cited generation · Streamlit on Hugging Face Spaces

Led technical delivery of hybrid retrieval with cited answers over library collections: MRR 0.049 to 0.209 and Hit@10 0.029 to 0.429 on 35 held-out queries.

2026

Generative Augmentation for Hand Osteoarthritis

CycleGAN · WGAN-GP · conditional VAE · ResNet-18 · 41,051 radiographs

Tested three generators at five augmentation ratios under a 135:1 class imbalance; none improved KL4 recall.

2026

Wearable IMU Exercise Classification

PyTorch · CNN-Transformer · 164 Apple Watch recordings

Benchmarked five architectures with recording-level cross-validation; CNN-Transformer reached 90.0% ± 5.7% accuracy.

Publications

2026

Detecting Deepfakes for the Courtroom: A Multi-Modal Forensic Output Framework

Zulal Akarsu, Maryan Rizinski, Ming Zhang, Kyung-Shick Choi, Deven Shah, Ittoop Shinu Shibu, Chan Woo Shin, Lou Chitkushev. International Conference on Intelligent Digital Forensics and Cybersecurity (IDFC 2026). Accepted.

Awards

2026

Winner, Civic Hacks 2026

MLH Best Use of Gemini API and infrastructure track, for ForTheCity, sidewalk accessibility validation

2025

4th place, BU DS+X Hackathon

For CollabNet, research collaborator discovery over OpenAlex

2024

Singtel Spot Award

Insurance Business

2023

Singtel Star Award

Customer Success

Skills

Python · PyTorch · scikit-learn · LightGBM · SQL · Spark · Databricks · Snowflake · MLflow · Airflow · Docker · RAG · pgvector · Neo4j · LangGraph