I build AI agents that have to work inside someone else’s real system, not in a demo.
Forward deployed engineer at Checksum AI, where I work on the agent that generates and heals end to end tests, the CLI and MCP server it runs through, and the evals that prove it actually got better.
Work
Forward Deployed Engineer · Checksum AI
I own the agent that generates and heals end to end Playwright tests for enterprise customers, plus the CLI and MCP server it ships through. Built a 400 case eval benchmark and RL environment from real production failures, then rebuilt the agent’s context pipeline against it.
eval 54% → 85% · 1,200+ tests per deploy · manual QA down 72% · 20+ enterprise customers
AI Software Engineer · Hyperlink, founding team
Early engineer on a browser agent that carries out tasks written in plain English. Shipped a Chrome extension to 400+ monthly users as sole engineer, rebuilt the AI video pipeline by profiling before optimizing, and made GPU vendor onboarding self serve.
latency 11s → 0.7s · 150 vendors, 580 GPUs onboarded in 40 min · scripting work down 65%
Software Engineer · FIS
Backend and data engineering for banking clients, mostly the old systems nobody wanted to touch. Traced weekend long migrations to full table scans with EXPLAIN plans and rewrote row by row logic into bulk operations.
migrations 87% faster · weekend runs cut to overnight
A native IDE for running several AI coding agents side by side, written in Rust on Tauri. Multiple terminals, a plugin API, and none of the Electron weight.
~300MB of memory where Electron takes 2 to 3GB · Rust · Tauri · TypeScript
I came to the US as an international student with no network and found my first opportunity through a university job board, so I have a lot of time for anyone doing that right now.
MS in Computer Science from UC Davis. Before Checksum I was at Hyperlink, and before that FIS, which is where I learned the hard part is usually the old system nobody wants to touch. On the side I build tools I actually use.
Contact
Happy to talk about agents, evals, or getting models to behave in production.