Enterprise RAG Platform
Document-intelligence platform with intent-driven retrieval, a feedback-driven evaluation harness, and citation-grounded answers that refuse unsupported claims.
PRODUCTION AI · AGENTIC SYSTEMS · ENGINEERING LEADER
RAG, conversational analytics & MCP-backed multi-agent systems with guardrails, grounding, human approval, and real evals — backed by 10+ years building distributed systems, including at Google, Uber & Microsoft, now as a hands-on CTO.
About
I'm an engineering leader and hands-on builder. After a decade building large-scale systems — including at Google, Uber, and Microsoft — I now design production AI and agentic systems as a startup CTO, leading teams where LLMs and agents are teammates that help us build fast and reliably.
My focus is AI that holds up in production: retrieval and agent systems with deterministic guardrails, citation grounding, human-in-the-loop approval, and real evaluation — not demos. Deep CS fundamentals and architecture judgment underpin everything I ship.
How I build AI that holds up
Risky actions pass through deterministic guards and human approval the model can't bypass — a SQL validator that rejects unsafe queries by construction, writes that execute only against a single-use, signed approval. “Be careful” is not a control.
Every response traces to a source and is allowed to say “I don't know” — citation-grounded retrieval with strict-evidence refusal, answers labeled authoritative, derived, or unverified.
Agents get golden sets and regression evals — routing, tool-choice, result-equivalence — so a prompt change is provable, not a vibe.
A decade of distributed-systems and data-infra work is what keeps these up under real load. The model is one component; the boundaries, state, evals, and observability around it are what make it production software.
Selected work
Three production systems built for one enterprise ERP client, sharing one safety and evaluation practice: Operations Copilot, NL2SQL Data Agent, and Knowledge Assistant.
Document-intelligence platform with intent-driven retrieval, a feedback-driven evaluation harness, and citation-grounded answers that refuse unsupported claims.
OLAP analytics agent that turns business questions into guarded SQL behind a deterministic SQLGlot safety layer, with a semantic metadata layer and bounded repair.
Agentic operations over business entities through governed MCP tools, with specialist routing and a human-in-the-loop approval boundary for risky writes.
Experience
Education
Lab
Self-directed prototypes exploring multimodal RAG, multi-agent systems, and classical ML.
Multimodal RAG with hybrid retrieval (dense + BM25), a LangGraph quality loop, RAGAS evaluation, and quality-gated human approval.
A supervisor-pattern multi-agent customer-service system (flights, hotels, car rental) with business-rule validation and session memory.
Competition-grade regression: gradient-boosting ensemble evolving to a multi-seed MLP, with rigorous feature engineering.
Skills
Contact
Open to AI/agentic engineering and staff-level roles, and consulting engagements.