Akshath Tiwari

Projects

A selection of LLM fine-tuning, agent, and evaluation systems, from production services to first-principles foundations.

Showing 9 projects

LLM Fine-Tuning & Model Optimization

Fine-Tuning Small LLMs to Match GPT-4.1 on Candidate–Job Matching

Fine-tuned Qwen3 (0.6B/4B, quantized) to reproduce a GPT-4.1-quality FitScore engine, closing a blinded A/B win-rate gap to within 1.9 points of the teacher model while running far cheaper.

UnslothHuggingFace TRLPEFT (QLoRA/LoRA)GRPO

Production LLM Services

Production LLM FitScore Microservice

FastAPI service serving a fine-tuned LLM on vLLM for real-time job–candidate scoring, with structured output, skill-gap analysis, and automatic OpenAI fallback.

FastAPIPydantic v2vLLM (xgrammar)MongoDB

Prompt Engineering & Optimization

GEPA-DSPy Prompt Optimization for Profile Matching

Applied reflective/evolutionary prompt optimization (GEPA via DSPy) to a candidate profile dedup and merge-decision system, cutting prompt development time from weeks to hours.

DSPyGEPAOptunaOpenRouter

Data Labeling & Evaluation

Golden Dataset Pipeline: LLM-as-Judge + Argilla Annotation

6-stage pipeline turning raw scoring outputs into a validated golden dataset, combining LLM-as-judge pre-labeling with human review in Argilla.

ArgillaFastAPIReact/Vite/TSKafka

AI Agents

LangGraph Deep-Research & Hiring Intelligence Agent

A LangGraph state-machine agent that plans, searches, writes, and grades its own research — applied to automated hiring-intelligence reports for enterprise clients.

LangGraphLangChainTavilyEXA

AI Agents

Contributing to LangChain's open_deep_research

Ran and contributed to LangChain's flagship open-source deep research agent — a supervisor + sub-researcher multi-agent system that placed #6 on Deep Research Bench.

LangGraphLangChainTavilyMCP

Load Testing, Benchmarking & Infra

vLLM Inference Benchmarking & Load Testing

Capacity-planning framework benchmarking vLLM serving performance directly at the engine level, plus a Locust-based load tester that measures accuracy and throughput simultaneously.

vLLMAWS SageMakerLocustgevent

Foundational / Educational

Building a GPT-Style LLM From Scratch

Implemented a GPT-style language model end-to-end from first principles — tokenization, BPE, embeddings, attention — following Sebastian Raschka's methodology, to build real intuition beneath the frameworks.

PyTorchtiktoken

Foundational / Educational

Positional Embeddings Deep-Dive: From Sinusoidal PE to RoPE

A numerical and visual investigation of why self-attention needs positional information at all, and why absolute positional encodings eventually break down — motivating RoPE.

NumPyPyTorchMatplotlib