ML systems engineer specializing in GPU-accelerated inference and distributed AI infrastructure — currently a Staff ML Engineer at Visa, and a contributor to vLLM's production stack. I like understanding systems from first principles, from transistors to transformer inference.
A study guide for SQL interview prep, distilled into a beautiful zero-dependency HTML page you can study anywhere.
I'm a software engineer who started with blinking LEDs and ended up building ML infrastructure at scale. The journey from soldering 8051 microcontrollers to deploying GPU workloads on Kubernetes has been anything but linear — and that's what makes it fun.
My background spans embedded systems, Linux kernel development, real-time operating systems, and now machine learning platforms. I believe the best engineers are the ones who understand the full stack — from transistors to transformers.
Lately that's meant going deeper into LLM inference internals — KV cache, quantization, speculative decoding — and GPU communication primitives like NCCL, rather than stopping at the framework API. When I'm not writing code, I'm writing about it — these courses are my way of giving back.
Architecting on-premise and cloud ML training infrastructure and building agentic AI systems at scale.
Merged a priority-routing feature into vLLM's production-stack, plus doc and benchmark contributions to vLLM itself — open-source LLM serving infrastructure.
Contributing to RTEMS RTOS — GPIO, PWM drivers, and I2C drivers for BeagleBone Black. Student & mentor.
From Linux kernel patches to embedded NTP clients for protection relays — I love the low-level stuff.
A mix of hardware and software — from microcontroller boards to open-source operating systems.
Merged feature contribution adding priority-based request routing to vLLM's production-stack, an open-source LLM serving infrastructure project.
Built GPIO test code, PWM driver, and I2C driver for BeagleBone Black within the RTEMS real-time OS. Mentored by Worth Burruss and Dr. Joel.
Implemented network time sync for protection relays at Easun Reyrolle — replacing costly coaxial cables with LAN-based NTP, reducing per-unit costs.
AI-powered wiki generator that parses codebases using tree-sitter, builds dependency graphs, and generates documentation with LLMs.
Open-source macOS utility that blocks every keystroke system-wide behind a full-screen overlay, so you can clean your keyboard without launching apps, retyping documents, or waking the machine.
Live-synced from GitHub — merged pull requests to projects I don't maintain.
Find me on the internet — always happy to chat about engineering, ML, or open source.