PRODUCTION AI · AGENTIC SYSTEMS · ENGINEERING LEADER

I build production AI that holds up

RAG, conversational analytics & MCP-backed multi-agent systems with guardrails, grounding, human approval, and real evals — backed by 10+ years building distributed systems, including at Google, Uber & Microsoft, now as a hands-on CTO.

GoogleUberMicrosoft
builder-CTO @ Yuanbao
10+ yrs10–30K QPS at ~140ms p9999.97% backup health
LangGraphRAGMCPGoJavaPythonGoogle Cloud
Wuhan, China felixhuhao@gmail.com

About

Leader who builds.

I'm an engineering leader and hands-on builder. After a decade building large-scale systems — including at Google, Uber, and Microsoft — I now design production AI and agentic systems as a startup CTO, leading teams where LLMs and agents are teammates that help us build fast and reliably.

My focus is AI that holds up in production: retrieval and agent systems with deterministic guardrails, citation grounding, human-in-the-loop approval, and real evaluation — not demos. Deep CS fundamentals and architecture judgment underpin everything I ship.

How I build AI that holds up

Safety lives at the boundary, not in the prompt.

Risky actions pass through deterministic guards and human approval the model can't bypass — a SQL validator that rejects unsafe queries by construction, writes that execute only against a single-use, signed approval. “Be careful” is not a control.

An answer without grounding is a guess.

Every response traces to a source and is allowed to say “I don't know” — citation-grounded retrieval with strict-evidence refusal, answers labeled authoritative, derived, or unverified.

If it isn't measured, it isn't shipped.

Agents get golden sets and regression evals — routing, tool-choice, result-equivalence — so a prompt change is provable, not a vibe.

The hard part is the systems, not the model.

A decade of distributed-systems and data-infra work is what keeps these up under real load. The model is one component; the boundaries, state, evals, and observability around it are what make it production software.

Selected work

One ERP AI layer: knowledge, analytics, operations.

Three production systems built for one enterprise ERP client, sharing one safety and evaluation practice: Operations Copilot, NL2SQL Data Agent, and Knowledge Assistant.

Enterprise RAG preview
Knowledge

Enterprise RAG Platform

Document-intelligence platform with intent-driven retrieval, a feedback-driven evaluation harness, and citation-grounded answers that refuse unsupported claims.

Vue 3FastAPILangGraphMilvus
NL2SQL Agent preview
Analytics

NL2SQL Data Agent

OLAP analytics agent that turns business questions into guarded SQL behind a deterministic SQLGlot safety layer, with a semantic metadata layer and bounded repair.

Vue 3FastAPISQLGlotQdrant
Ops Copilot preview
Operations

ERP Operations Copilot

Agentic operations over business entities through governed MCP tools, with specialist routing and a human-in-the-loop approval boundary for risky writes.

DeepAgentsFastAPISpring AI MCPJava 21

Experience

10+ years, top to bottom of the stack.

2025 — Present
Chief Technology Officer · Yuanbao Creative Technology
Architected a unified, agentic AI layer for an enterprise ERP SaaS platform — operations, analytics, and knowledge — leading a team of 4 (see projects above).
2023 — 2025
Independent Software Engineer & Consultant · Independent
Shipped a cloud-native travel-guide SaaS on AWS and a portfolio of RAG/agent prototypes — bridging hands-on practice into the CTO role.
2022 — 2023
Senior Software Engineer · Uber
Built Uber's first CPG Ads backbone (10–30K QPS at ~140 ms p99) and the GTIN-keyed multi-location campaign model; led the phased production rollout.
GogRPCDocStoreKafkaFlinkPrometheus/Grafana
2018 — 2021
Software Engineer · Google
Cloud SQL: enabled reliable point-in-time recovery and raised backup health from 99.92% to 99.97%; ran data-plane on-call to a 99.95% SLA.
JavaGoMySQL HACSV import/export GACloud SQL IAM
2014 — 2017
Software Development Engineer · Microsoft
Exchange/Outlook: built communication-signal and migration tooling supporting 400M+ Hotmail users moving to Outlook.
C#ExchangeGraph People APISQLMonitoring/Alerting
2011 — 2013
Software Developer / Data Engineer · Epic Systems · HCR ManorCare
Release-automation tooling, plus SSIS ETL pipelines and advanced SSRS reporting.

Education

Where the fundamentals come from.

2010 — 2012
Bowling Green State University · M.S. in Computer Science
GPA 4.0/4.0 · Bowling Green, OH
2006 — 2010
Wuhan University · B.S. in Software Engineering
GPA 3.5/4.0 · Wuhan, China

Lab

Personal R&D.

Self-directed prototypes exploring multimodal RAG, multi-agent systems, and classical ML.

Multimodal RAG

Multimodal RAG with hybrid retrieval (dense + BM25), a LangGraph quality loop, RAGAS evaluation, and quality-gated human approval.

MultimodalRAGASLangGraphMilvus

Multi-Agent Travel Assistant

A supervisor-pattern multi-agent customer-service system (flights, hotels, car rental) with business-rule validation and session memory.

Multi-agentSupervisorLangGraph

Vehicle Price Prediction

Competition-grade regression: gradient-boosting ensemble evolving to a multi-seed MLP, with rigorous feature engineering.

PyTorchGradient boostingFeature engineering

Skills

Technical toolkit.

LLM & Agents

Agent orchestrationAgent governanceTool / function callingMCP integrationMulti-agent routingContext engineeringLangChainLangGraphDeepAgentsOpenAI / Claude / DeepSeek

RAG & Retrieval

RAG (multimodal)Hybrid retrievalEmbeddingsRerankingQuery optimizationMilvusQdrant

Evaluation & Safety

LangSmithRAGASOpenEvalsLLM-as-judgeGuardrailsHuman-in-the-loopObservability

Programming Languages

Java & Go (Google readability)PythonSQLC#C++

Backend & Distributed Systems

MicroservicesAPI designgRPC / ProtobufRESTThriftSpring AIFastAPISSE

Data Systems

Data modelingPostgreSQLMySQLRedisBigQueryClickHouseKafkaFlinkMongoDB

Cloud & Platform

Google CloudAWSDockerKubernetesLinuxCI/CD

ML / Data Science

PyTorchCatBoost / LightGBM / XGBoostFeature engineeringTime-series

Contact

Let's build something reliable.

Open to AI/agentic engineering and staff-level roles, and consulting engagements.