Hi, I’m Mark
I'm an ML Engineer at 314e Corp in Bengaluru, where I build and operate MLOps pipelines for healthcare ML systems — orchestration, distributed data processing, and model serving at scale. Outside of work I write low-level Rust and Python for fun: reimplementing NumPy from scratch, and building terminal tools I actually use every day.
- Built and owned the MLOps pipeline for a clinical patient entity-extraction model serving multiple healthcare clients, raising accuracy from 72% to 94% on an internal held-out set via a human-in-the-loop retraining loop.
- Orchestrated fault-tolerant training and serving workflows on Temporal and SkyPilot across RunPod and Hugging Face Endpoints, standardizing deploys across models.
- Benchmarked open Vision-Language Models (incl. Gemma 3) with a reproducible ClearML test suite, surfacing a Flash-Attention/SDPA correctness issue in the serving path.
- Cut model cold-start latency 77% by pre-packaging weights and tokenizer into the runtime image and warming the serving process.
- Cut model training time from 90 hours to 25 hours on >500GB datasets via targeted hyperparameter tuning, and eliminated OOM errors with a custom memory-efficient iterable dataloader.
- Re-engineered single-node Python ETL into distributed PySpark jobs over FHIR resources for multiple clients, cutting data-loading time ~80% and unblocking training on 500GB+ datasets.
Engineering writeups of this platform, which I contributed to — published by 314e, authored by Dr. Srivatsan Sridhar: AI document extraction & classification · IDP for medical records
- Engineered and optimized high-frequency trading strategies, reducing data-processing latency through targeted performance tuning.
- Built and deployed ML models from scratch against live market APIs to execute hedging strategies, backed by a comprehensive unittest suite for core trading libraries.
- Built a data-preprocessing pipeline for semantic-segmentation models over large-scale GIS raster data, and implemented distributed computer-vision algorithms (Douglas–Peucker, Jarvis March).
- Deployed scalable Django apps on AWS with an end-to-end CI/CD pipeline, and hardened a high-performance raster-processing app through refactoring and unit tests.
- Applied NLP techniques to proprietary Amazon-review datasets to deliver product-trend insights and predictive models; automated preprocessing in Python/Bash, cutting data redundancy by 90%.
ikj loop order against BLAS-backed NumPy. Spec-driven with ~30 milestone tests.mpv over its JSON IPC socket, and scrobbles to Last.fm via cmusfm. Vim-style keys, live search, play history, and pixelated album art rendered straight in the terminal.More on GitHub →
B.Tech, Computer Science — ABV-IIITM Gwalior, India · SGPA 8.67 · 2024
Relevant coursework: Artificial Intelligence, Statistics, Cloud Computing
DevOps on AWS Specialization · GCP Network Deployment · Open-source contributor @PyMC
