I go broad and I go deep: shipping full-stack products and advising startups on one side, optimizing GPU kernels and maintaining a large-scale ML systems codebase on the other.
Currently: maintaining CS249r (Harvard's Machine Learning Systems course repo), Triton/CUDA kernel optimization, and full-stack + AI product work.
- Maintainer, CS249r · Machine Learning Systems (Harvard), reviewing and merging PRs, triaging issues, and hardening CI across the TinyTorch, StaffML, MLSys·im, and Labs sub-projects.
- Kernel optimization work in Triton/CUDA: GEMM tuning, memory coalescing, occupancy and tiling experiments.
- Full-stack product and AI consulting work with startups and product teams.
- IIT Guwahati, Class of 2028. Kaggle Grandmaster.
I don't fit neatly into "generalist" or "specialist," I'm both, depending on what the problem needs. Some days that means tuning a CUDA kernel for warp-level occupancy. Other days it means sitting with a founder to figure out which parts of their AI product actually need to be built versus bought, then shipping the MVP myself.
That range comes from genuinely enjoying both ends: practical business problems that need clean, fast execution, and deep infrastructure problems that need low-level systems thinking. I've worked across early-stage startups, hackathon teams, research-oriented engineering groups, and enterprise workflows, and I'm comfortable being the person who bridges a technical team and a non-technical one when a project needs that translation.
Underneath all of it is the same curiosity: understanding systems from the inside out, whether that system is a GPU kernel's memory access pattern, a startup's infrastructure spend, or a course repository's CI pipeline.
Harvard's open-source ML Systems Engineering course repository. I maintain the repo across its sub-projects:
- TinyTorch: a from-scratch deep learning framework built to teach tensor abstractions, autograd, computational graphs, and backend execution.
- StaffML: an interview-prep question bank and practice app, physics-grounded ML systems questions with a Cloudflare Workers backend.
- MLSys·im: a first-principles analytical modeling framework for ML systems, also the physics engine behind the browser-based interactive labs.
- Labs: 34 browser-based, WASM-exported interactive labs built on MLSys·im.
Maintainer work includes reviewing and merging contributor PRs, root-causing CI failures, fixing silent data-loss and security bugs, and writing contributor-facing system design documentation for the ecosystem.
Beyond my own projects, I work with startups and product teams on:
- choosing efficient AI/ML architectures
- optimizing infrastructure costs
- scaling products pragmatically
- improving engineering workflows
- balancing performance with maintainability
- shipping faster without sacrificing qualitylanguages python, c++, cuda, javascript, typescript, c, R
ml/ai pytorch, triton, tensorflow, jax
systems cuda, distributed systems, gpu programming
backend node.js, express, fastapi, svelte, sveltekit
frontend react, next.js
infra linux, docker, git, vercel, kubernetes- Portfolio: shashankt.vercel.app
- LinkedIn: linkedin.com/in/rocky0714
building systems that make AI workloads faster, scalable, and usable in the real world.




