Lead AI Engineer
Jan 2024 — PresentAPS Asia PDA Enterprise · Rockville, MD
- Designed and shipped production GenAI and multi-agent systems with Python, PyTorch, LangChain/LangGraph, AWS Bedrock, MCP, A2A, FastAPI, and RAG architectures, improving LLM application reliability and response quality by 28%.
- Architected high-throughput LLM inference on AWS EKS with vLLM, KV caching, quantization, FlashAttention, continuous batching, and speculative decoding, cutting inference latency and serving costs by 68%.
- Fine-tuned transformer LLMs with LoRA/PEFT adapters on domain-specific datasets, improving task accuracy by 15% while reducing fine-tuning compute by 70%.
- Hardened LLM applications with prompt-injection detection, input/output sanitization, guardrails, PII filtering, and tool-access policies, reducing unsafe or unauthorized model interactions by 30%.
- Built MLOps/LLMOps pipelines covering CI/CD for model and prompt versioning, automated evaluation, and OpenTelemetry with LangSmith tracing, improving production observability and debugging efficiency by 25%.
- Designed semantic-search infrastructure on embedding models, FAISS, ChromaDB, and OpenSearch for low-latency retrieval over large unstructured datasets.