Achieved ~99% code evaluation accuracy through AI-powered logical validation without relying on traditional compilers
Enabled multi-language support (Python, Java, C++, JavaScript) via intelligent pattern-based detection and cross-language sanity checks
Delivered real-time curriculum-linked progress automation through seamless mapping of submissions to TodoItems, Topics, and Subroadmaps
Generated structured JSON verdicts enriched with human-like explanations, corrected code suggestions, and complexity analysis
Reduced evaluation inconsistencies with production-grade error classification including COMPILATION_ERROR, RUNTIME_ERROR, and VALIDATION_ERROR pipelines
Enhanced AI reasoning through RAG-based semantic context retrieval using text-embedding-ada-002 and reference solution augmentation
Ensured zero-downtime evaluation using multi-model fallback orchestration with o3-mini, o1, and gpt-35-turbo routing
Supported scalable dev/prod deployments via environment-aware API routing, secure JWT endpoints, and fault-tolerant backend flows
Architected an end-to-end AI-powered code evaluation system replacing traditional compilers with RAG-enhanced logical judgment, leveraging semantic retrieval, model-context engineering, and multi-model orchestration to achieve 99% evaluation accuracy across Python, Java, C++, and JavaScript.
Built a multi-stage language detection engine using regex patterns, anti-pattern suppression, syntax heuristics, and confidence-based classification to prevent cross-language submissions and ensure evaluation integrity for every code block.
Implemented a production-grade MCP-compliant prompt pipeline generating strictly structured system/user message arrays, including judge instructions, evaluation rules, test-case schemas, complexity requirements, and JSON-first verdict formatting.
Designed a dual-layer response parsing system with JSON block extraction, Markdown fallback resolution, regex-based error isolation, and verdict normalization to guarantee consistent outputs even with noisy AI responses.
Engineered a multi-model AI orchestration layer dynamically routing requests between o3-mini (accuracy), o1 (reasoning), and gpt-35-turbo (performance) with token-window optimization and context-aware selection.
Integrated a RAG pipeline with ChromaDB using text-embedding-ada-002 to retrieve reference solutions, constraints, edge cases, and complexity hints, enabling AI to perform context-enriched evaluation rather than plain code matching.