Scientific ML & engineering surrogates
CFD field surrogates, graph neural networks, uncertainty quantification, simulation-level generalization, physics checks, and design optimization.
Georgia Tech · Ex-Rolls-Royce
Machine learning for engineering, data, and real-world decisions.
I'm a Machine Learning Engineer with experience in scientific ML, model evaluation, retrieval, time-series data, and ML systems. My work started in aerospace, which taught me to care about the data, the limits of a model, and whether people can actually use it.
I've worked on scientific ML, model evaluation, retrieval, time-series data, and ML systems. I like clear questions, solid tests, and explaining results in a way that engineers and non-technical teammates can use.
CFD field surrogates, graph neural networks, uncertainty quantification, simulation-level generalization, physics checks, and design optimization.
Hybrid retrieval, reranking, recommender systems, evidence-grounded RAG, LLM/VLM evaluation, protected tests, provenance, and failure analysis.
FastAPI, Docker, Kubernetes, Kafka, PyTorch distributed training, ONNX, Core ML, Qualcomm QNN, Android/JNI, observability, and deployment-focused benchmarking.
NewsLens: chronological recommendation evaluation and model release controls. AeroSynth-Eval: human-aligned evaluation and controlled augmentation experiments.
AIRFAANS: geometry-aware CFD surrogates. Surrogate Model Learning: grouped regression, uncertainty diagnostics and measurement requests.
EdgeGenBench: model export and iOS inference instrumentation. Equity Backtest: temporal evaluation, accounting tests and data-quality controls.
Georgia Tech AE 6394 coursework comparing a pointwise MLP, MeshGraphNet-style GNN, and point neural operator on official AirfRANS meshes with simulation-level splits, force verification, and resumable full-mesh evaluation.
Across 3 matched seeds × 3 architectures × 200 official interpolation test meshes, MeshGraphNet had the lowest mean error for all predicted fields and drag, while the point operator had the lowest mean lift error. Reynolds/AoA OOD, uncertainty, and active-learning studies are now in progress; no results are reported until the matched runs finish, so the project is not presented as operationally ready.
Public engineering-data studies of Gaussian-process and conventional surrogate models using grouped splits, multi-seed analysis, split-conformal intervals, and a distance-to-training-domain guard.
An airfoil Gaussian process achieved R² 0.8145 on a physically grouped split, while the 10-seed mean was 0.8662 ± 0.0680. Nominal 90% intervals did not retain 90% coverage under design shift—an explicit result that shaped the uncertainty analysis.
A moderation benchmark and fail-closed service with versioned policy controls, a checksummed model registry, human-review routing, shadow comparison, rollback, telemetry, and an AWS infrastructure plan.
On 97,320 held-out public Civil Comments, validation-selected thresholds reduced false acceptance from 11.43% to 1.84%. The frozen three-way candidate then falsely allowed 59.32% of 2,802 human-annotated ToxicChat prompts. A BeaverTails-only binary experiment lowered false acceptance to 18.79%, but raised false rejection to 16.43% and cannot escalate, so it is not a replacement.
Evaluation-first GenAI system over 3,233 citation-preserving NASA NTRS chunks, combining BM25 + dense retrieval, Reciprocal Rank Fusion, cross-encoder reranking, pgvector, evidence-sufficiency checks, and application-controlled citations.
The public evidence track now combines QASPER and SciFact retrieval, an audit of 20,283 TREC relevance judgments and 2,840 citation-support judgments, and a controlled NASA test where all 200 wrong source IDs were rejected. The 50-case aerospace author audit is still pending.
Leakage-safe MIND recommendation paired with a separate Go, Kafka, PostgreSQL, and FastAPI article path built for idempotency, dead letters, and freshness-aware ranking.
NDCG@10 0.3664 offline. In a 500-event local run, publish p99 was 44 ms, index-freshness p95 was 79 ms, all duplicate probes were caught, and a stopped consumer's partition recovered in 5.7 seconds.
A real NASA DASHlink flight-anomaly track plus a separately labeled generated aircraft-design benchmark spanning ONNX, Core ML, Qualcomm QNN, native C++, Android, and browser inference.
The recorded-flight model reached 0.7380 macro F1 on 17,780 aircraft-disjoint approaches and remained blocked by its macro-F1 and late-flap-recall gates. ONNX prediction consistency stayed above 99.55% under the tested sensor corruptions.
Aircraft inspection-image evaluation using the public AGDD real-image benchmark, with generated images kept as explicit controls or augmentation rather than treated as operational evidence.
Across ten matched seeds, mixed training raised mean macro F1 from 0.3881 to 0.4419 but lowered mean crack recall from 0.4000 to 0.3167. A separate 3,224-image DLR aircraft-dent track reached 0.9777 dent recall on its first 645-image test, but only 0.5969 ROC-AUC because of false alarms. It remains a baseline, not a maintenance claim.
Built a leakage-resistant failure-detection workflow over 15,120 simulated thrust profiles using time-series feature extraction, statistical analysis, clustering, and supervised classification.
96.8% accuracy · 0.986 ROC AUC on previously unseen motor geometries, with emphasis on generalization and failure analysis.
Other work and active studies: Atlanta Mobility Resilience Digital Twin · GREEN TEA systems-of-systems modeling · walk-forward equity backtesting
Graduate Research Assistant under Prof. Dimitri Mavris · Machine Learning & Applied AI
Built time-series ML for rocket-motor failure detection across 15,120 simulations; contributed to Delta Air Lines-sponsored HERO work on source-aware retrieval for safety analysis; and developed surrogate models for sustainable-aviation studies.
Machine Learning Engineer
Built Python-based diagnostic, anomaly-detection, predictive-maintenance, and 1 Hz/10 Hz time-series workflows for multivariate aircraft-engine sensor data using Azure Databricks, working with lifecycle engineers to validate failure modes and maintenance insights.
Data Science Intern
Worked on applied ML for aircraft-engine performance and component health, including exploratory data analysis, feature engineering, predictive modeling, anomaly detection, and model evaluation.
Student Chair
Led graduate-student communications and community engagement, coordinating peer networking and chapter outreach.
Student Volunteer · SRM University
Organising Team Member · SRM University
Team Participant · SRM University
Worked in a five-member team on a sustainable urban-transport concept for densely populated Indian cities, comparing travel time, fuel use, cost, and CO2 emissions.
I'm looking for Machine Learning Engineer and Applied ML roles.
For Machine Learning Engineer, Applied ML, Retrieval, or ML Evaluation opportunities, contact me at tsarkar34@gatech.edu.