Served as the flagship chat surface of a platform tested with 1000+ active users, processing 18,000+ AI queries from 12,000+ unique visitors and gathering 800+ detailed user feedback that drove 3 major product iterations to a 99% user satisfaction rate - all with no API keys required from users.
Achieved ~99% RAG accuracy and 95% hallucination reduction through the 3-Layer Hybrid RAG pipeline over 10M+ ChromaDB embeddings - Layers 1 and 2 (13 regex intent rules + 12 pre-cached semantic anchors) resolve most queries with zero additional embedding API calls, and full context retrieval completes in <1 second, so every answer is grounded in verifiable production facts instead of generic LLM filler.
Sustained sub-300ms first-token latency across all 24+ models and 30+ streaming endpoints via dual-channel delivery - when a provider degrades, traffic reroutes between Azure AI Foundry and AWS Bedrock with zero downtime, so a single provider outage never takes ARX offline.
Improved perceived response latency by 60% through adaptive SSE streaming (2–8 char chunks, ±3ms jitter, <50ms intervals) combined with the ChromaDB response memory cache - repeated questions replay instantly from cache while feeling identical to live inference, cutting redundant LLM token spend with zero buffering under high load.
Eliminated model-release lag entirely: the MongoDB ActiveAIModels registry means new frontier models (GPT-5.4, Claude Opus 4.8, Grok 4.3) go live to users within minutes of provider release by flipping a database record - zero code changes, zero redeploys - keeping ARX permanently current in a market where model generations turn over monthly.
Reduced API abuse by 99% on infrastructure handling 50K+ daily API calls: every chat request clears JWT auth and Redis-backed rate limiting before any model call is made, so not a single unauthorized request has ever triggered LLM spend - with structured 429 envelopes rendering accurate live countdown timers without any frontend polling.
Kept answers permanently fresh with zero manual content operations: the daily cron re-ingestion pipeline re-embeds all 12 data sections from MongoDB into ChromaDB automatically, so profile, project, and experience updates propagate into RAG answers within 24 hours without anyone touching the pipeline.
Built on the same production AI infrastructure powering systems that reached 600K+ users, drive 80%+ of platform traffic, and run a <40ms AI inference review engine - ARX is a customer-facing window into battle-tested architecture, not a demo assembled for screenshots.
Generated direct inbound opportunity: multiple job calls and offers from AI companies, and 3–4 founders reached out to build similar multi-model chat products after using ARX and seeing the architecture respond in real time - the product itself became the strongest proof of engineering capability.
Made unauthorized model spend structurally impossible rather than merely unlikely: because JWT verification and quota checks both resolve before the controller is entered, there is no code path in which an unauthenticated or over-quota request reaches a provider API - and the 4-layer perimeter rejects a direct curl for a missing Origin header before authentication is even attempted, closing the exact gap that CORS alone structurally cannot cover.
Cut retrieval cost to zero LLM calls per query by moving query expansion into the corpus at index time instead of paying for query rewriting on every request - where a conventional RAG pipeline spends multiple LLM calls per query on rewriting and HyDE, this one spends none, and intent routing resolves the most common questions with no embedding call at all, turning them into pure metadata lookups while the remaining layers absorb everything else.
Architected ARX (AI Architect) - a standalone, product-grade conversational AI application at /arx built on Next.js 15 (App Router), React 19, TypeScript 5, and Tailwind CSS 4 - giving direct chat access to 24+ frontier models across 11 providers including OpenAI (GPT-5.4 Pro, GPT-5.4 Mini, GPT-5.3 Codex), Anthropic (Claude Opus 4.8, Opus 4.7, Sonnet 4.6), Google (Gemini 3.1 Pro, Gemini 3 Flash, Gemma 3 27B), xAI (Grok 4.3, Grok 4.2 Reasoning), DeepSeek (V4 Pro, V4 Flash), Meta (LLaMA 4 Maverick), Mistral (Mistral Large 3), Moonshot (Kimi K2.6, K2.5), Qwen (Qwen3 VL), Nvidia (Nemotron Nano), and MiniMax (MiniMax M2) - all served through 30+ streaming endpoints with intelligent model routing and sub-300ms first-token latency.
Engineered the full request pipeline surfaced to users as a 6-stage animated workflow: Your Query → JWT Auth & Rate-Limit → RAG Retrieval (ChromaDB vector search pulling full personal context in <1s) → Model Router (picks the best of 24+ frontier models) → AI Inference → Token Streaming - every stage backed by the real production backend: Passport.js JWT validation, Redis-backed rate limiting with structured 429 rateLimit envelopes (limit, remaining, resetInSeconds) that the frontend converts into live countdown timers without polling, and SSE token streaming rendered incrementally in the chat UI.
Grounded every answer through a 3-Layer Hybrid RAG pipeline over 10M+ ChromaDB embeddings (Azure OpenAI text-embedding-3-large on self-hosted Azure VMs): Layer 1 - 13 regex intent-detection rules for zero-latency section routing; Layer 2 - 12 pre-cached semantic anchor embeddings compared via cosine similarity with no extra API calls; Layer 3 - full-corpus similarity search fallback; combined with entity-diverse re-ranking (guarantees ≥1 document per company/project) and keyword priority boosts (critical=+3, high=+1) - achieving ~99% RAG accuracy and 95% hallucination reduction so ARX answers about real experience, real architecture decisions, and real production numbers instead of generic LLM output.
Built a dual-channel model delivery layer: Channel 1 routes through Azure AI Foundry (Anthropic Foundry SDK) and direct provider APIs (Azure OpenAI, Google GenAI, Mistral, DeepSeek, Grok, Llama, Kimi clients); Channel 2 routes through AWS Bedrock (ConverseStreamCommand) for Claude, Qwen, Gemma, MiniMax, Nemotron, and Kimi - with a MongoDB ActiveAIModels registry so models can be added, reordered, or disabled at runtime with zero redeploys; the frontend model selector hydrates from this registry via a Zustand aiModelsStore, letting users switch between any of the 24+ models mid-session without losing conversation context.
Delivered an adaptive token streaming architecture over Server-Sent Events with dynamic chunk sizing (2–8 characters per write) and ±3ms jitter timing to prevent synchronization artifacts across providers with wildly different native token cadences, sustaining <50ms chunk intervals uniformly across all 24+ models - complemented by a ChromaDB-backed response memory cache that re-streams previously computed answers chunk-by-chunk so cached responses feel identical to live inference, eliminating redundant LLM API calls for repeated questions and improving perceived latency by 60%.
Developed an MCP-compliant prompt engineering layer with dynamic system/user role injection and adaptive tone control (professional, casual, friendly, technical) so the same grounded facts render in the register the user asked for, plus Redis-backed long-term session memory preserving conversational context across messages and reconnects; coupled with a daily cron-driven RAG re-ingestion pipeline that automatically clears and re-embeds resume/project data from MongoDB into ChromaDB across 12 structured sections (salary, experience, projects, education, GitHub, mentorship, DSA, skills, achievements, social media, personal, major contributions) - ARX never answers from stale context with zero manual content operations.