Delivered a full production RAG pipeline in Go that fires up to 12 ChromaDB queries per user request (6 expanded query variants × 2 collections) with sub-200ms latency in fast search mode and 2-3 seconds end-to-end in full RAG mode - covering the complete path from query to streamed LLM answer.
Indexed 3,000+ LeetCode problems with full solutions in dsa-cluster and 3,000+ multi-platform DSA problems across 8 major competitive programming platforms in dsa-problem-list - ingested via a 10-worker concurrent Go pipeline with idempotent upserts that can safely re-run without duplicating data.
Implemented HyDE (Hypothetical Document Embedding) - a technique where the LLM generates a hypothetical solution paragraph that is embedded and used as a query vector, bridging the semantic gap between a short user query and a long solution chunk in embedding space, significantly improving solution retrieval recall over direct query embedding alone.
Applied Reciprocal Rank Fusion (RRF, k=60) following Cormack et al. (2009) to merge ranked result lists from multiple query variants - immune to embedding score scale differences across query texts - ensuring that documents consistently ranked high across many query variants are correctly promoted regardless of absolute similarity scores.
Built a problem-solution pairing system that joins problem and solution chunks by problem_id before MMR selection, so FinalScore is computed from the combined pair value - the LLM receives both the full problem statement and the solution explanation together in a single RAGContext, delivering richer instructional context than separate chunk retrieval.
Designed a dual difficulty field schema (difficulty_label string + difficulty_rating integer) enabling a single ChromaDB collection to serve both label-based platforms (LeetCode easy/medium/hard) and rating-based platforms (Codeforces 800–3500+) with the same $eq and $gte/$lte filter operators - eliminating the need for separate collections or platform-specific query paths.
Delivered graceful pipeline degradation: if GPT-5.2 query expansion fails, the pipeline continues with the original query alone; if HyDE generation fails, the pipeline proceeds with the 4 variants; if individual embedding calls fail during ingestion, the worker moves on without halting the run - all degradation paths are logged at warn level with structured fields for diagnosis.
Implemented an SSE streaming architecture with client disconnect detection via context.Canceled propagated through the streaming callback - ensuring GPT-5.2 streaming is cleanly terminated when the user closes the browser, preventing wasted token generation and goroutine leaks under concurrent load.
Built a MongoDB AI response cache (stored_responses collection) that eliminates redundant RAG pipeline executions and LLM token spend for repeated queries - on a cache hit, the stored response is replayed via a streamCached function that sends 80-rune chunks with 8ms intervals, making cached responses feel identical to a live GPT stream with no perceptible difference in UX; on a cache miss, the full response is collected via strings.Builder and persisted via a non-blocking goroutine upsert so the HTTP handler returns immediately; hitCount and lastHitAt are tracked per cache entry to surface the most frequently requested topics.
Shipped a production-grade authentication layer protecting all AI endpoints - combining email+OTP signup (Azure Communication Services / AWS SES), Google OAuth 2.0, and GitHub OAuth behind the same JWT middleware - ensuring no unauthenticated request can trigger a GPT-5.2 call and incur API cost; the OptionalAuth design on cognitive service routes allows guest access rate-limited by IP while still hydrating full user context for authenticated sessions.
Applied three-tier rate limiting with structured 429 responses directly controlling LLM token spend: the AI tier limits each user+IP to 3 requests per 5-hour window, directly capping worst-case GPT-5.2 cost per user to approximately 3 × 8,000 tokens = 24,000 tokens per window; the rateLimit.resetAt ISO timestamp in every 429 response lets the frontend render accurate countdown timers without polling; Redis-primary with in-memory fallback ensures rate limits hold even during Redis restarts.
Achieved full per-query execution traceability through structured zerolog checkpoint logs at every RAG pipeline step - every user query can be fully reconstructed from the log files alone, including which query variants were generated, which filters were applied, how many results were retrieved per collection, RRF scores, pairing results, and final MMR selection order.
Received interest from founders and engineers after sharing the RAG architecture - specifically the HyDE + RRF + Keyword Hybrid approach in Go, which is rare outside Python-based RAG frameworks - validating the architectural approach as production-worthy and reusable for other domain-specific AI tutoring systems.
Closed the gap between CORS and real access control by shipping a server-side Origin re-validation layer plus a shared reverse-proxy secret on every /api/v1/* route - meaning even a correctly-guessed API shape is unreachable by direct HTTP tooling (curl, Postman) without also possessing the internal proxy secret, independent of anything a browser would send.
Eliminated a class of false-logout incidents by design: the auth middleware distinguishes a transient database error (503) from an actually invalid session (401), so a MongoDB blip during peak load degrades to a retryable error for the frontend instead of silently ending every active user's session.
Stood up an operationally independent monitoring surface - a password-gated live status dashboard plus automated AWS SES restart alerts - giving real-time visibility into MongoDB/Redis/ChromaDB health and deployment events without any third-party observability vendor.
Reduced blast radius from datastore outages through asymmetric fail-fast/fail-soft startup design - MongoDB failures halt boot immediately (preventing the service from ever running in a broken auth/cache state), while ChromaDB and Redis failures degrade functionality instead of availability, so a vector-store or cache outage never takes down query answering entirely.
Architected a production-grade AI tutoring backend in Go built around a 6-step RAG pipeline: Step 1 (Query Expansion + HyDE) uses 2 GPT-5.2 calls to generate 4 semantic query variants plus a Hypothetical Document Embedding - a short solution paragraph written by the LLM that embeds close to solution chunks in vector space, dramatically improving solution retrieval recall; Step 2 (Multi-Vector Search) embeds all 6 queries using Azure OpenAI text-embedding-3-large and fires up to 12 ChromaDB queries (6 × 2 collections); Step 3 (RRF Merge, k=60) fuses all ranked lists using Reciprocal Rank Fusion, immune to embedding scale differences across query texts; Step 4 (Keyword Boost) adds deterministic signals (word-overlap ratio, exact prefix, per-word title/topic/platform matches) computing a HybridScore = RRFScore + boost; Step 5 (Problem-Solution Pairing) joins problem and solution chunks by problem_id into unified RAGContext structs; and Step 6 (MMR Diversity Filter, lambda=0.7) applies Maximal Marginal Relevance for a 70% relevance / 30% diversity balance, returning exactly topK results.
Engineered dual ChromaDB collections on self-hosted Azure and AWS VMs with runtime cloud switching via environment variable: dsa-cluster stores 3,000+ LeetCode problems as two-chunk documents per problem (problem-{id} containing full description, examples, constraints, and hints + solution-{id} containing the editorial) joined at retrieval time by problem_id - delivering both question and answer together to the LLM in a single enriched context; dsa-problem-list stores 3,000+ multi-platform DSA problems across 8 platforms: LeetCode, Codeforces, GeeksForGeeks, HackerRank, CSES, InterviewBit, HackerEarth, and CodeStudio - both collections using HNSW cosine-space indexing for length-normalized similarity across problems of varying description lengths.
Implemented platform-aware and difficulty-aware ChromaDB metadata filtering: detectPlatform() resolves 8 platform aliases from query text (lc/leetcode, cf/codeforces, gfg/geeksforgeeks, hr/hackerrank, cses, ib/interviewbit, he/hackerearth, cn/codestudio); detectDifficulty() handles both label-based difficulty (easy/medium/hard for LeetCode and GFG) and 15+ Codeforces numeric rating ranges (800 beginner through 3500+ grandmaster, each with a ±50 window) - producing ChromaDB $eq and $gte/$lte where filters composed with $and for multi-condition queries; and a dual difficulty field design (difficulty_label string + difficulty_rating int) that enables both label and numeric filtering on the same dsa-problem-list collection without schema duplication.
Built a multi-LLM orchestration layer integrating three distinct AI providers: GPT-5.2 (Azure OpenAI) as the primary model for query expansion, HyDE generation, and streaming answer delivery - correctly configured as a reasoning model using max_completion_tokens (not max_tokens) with no temperature or top_p parameters; Claude Sonnet 4.6 (AWS Bedrock) as secondary answer generator via Bedrock Runtime invoke endpoint; and Gemini 3.1 Pro (Google AI) as tertiary generator - with all three verified at server startup via TestAllModels() sending a test prompt and logging success/failure per provider before accepting any user traffic.
Designed a hybrid scoring system that outperforms pure vector search for exact DSA problem name queries without requiring BM25 or a separate sparse retrieval index: word-overlap ratio boost (0.0–2.0 based on matched significant words / total significant words), exact title prefix match (+1.5), per-word title match (+0.3 each), per-word topic match (+0.1 each), per-word platform match (+0.2 each), solution/problem type intent match (+0.3), explicit boost type match (+0.2), and a -2.0 penalty for empty titles - stacked on top of the RRF base score with all stop words and structural keywords excluded from word-overlap calculation.
Implemented a real-time SSE (Server-Sent Events) streaming architecture on Go/Gin delivering incremental GPT-5.2 response chunks with four typed event envelopes: status (pipeline progress), chunk (one LLM token/segment), done (stream complete with metadata), and error (terminal failure) - with streaming-specific HTTP headers (Cache-Control: no-cache, Connection: keep-alive, X-Accel-Buffering: no) injected per-route via Gin middleware and client disconnect detection via context.Canceled propagated through the streaming callback, enabling clean stream termination without goroutine leaks.