NPU · LIVE
INIT
booting inference core
LAYER ACT.0 tok/s
VJ/
AI / ML ENGINEER — LLM EVALUATION · AGENTIC SYSTEMS

VatsaJoshi

2+ years shipping production LLM systems — from eval harnesses and red-teaming to retrieval & memory infrastructure.

SCROLL
LANGGRAPH    PYTORCH    LORA / QLORA    DEEPSPEED    RAG    GRAPHRAG    QDRANT    FAISS    RED-TEAMING    FASTAPI    RUST    AWS    
( 01 )WHAT I DO

I build agentic AI systems, own LLM evaluation pipelines, and fine-tune models that hold up in production.

Legal AI, health assistants, retrieval & memory infrastructure — end to end.

0+
YEARS ENGINEERING
0x
LEGAL-NER EXTRACTION COST CUT
0%
DOC-LEVEL ZERO-FAILURE
0M
ROW GROUND-TRUTH EVAL SET
( 02 )FOCUS AREAS

Three things I go deep on.

01

Agents & Retrieval

Multi-agent orchestration, tool-calling & routing — plus RAG, GraphRAG and memory infrastructure owned end-to-end.

LANGGRAPH · CREWAI · GRAPHRAG · QDRANT
02

Fine-Tuning & Training

LoRA / QLoRA & DeepSpeed ZeRO distributed training — embeddings, rerankers and low-resource translation.

LORA / QLORA · UNSLOTH · DEEPSPEED
03

Evaluation & AI Security

Eval harnesses, golden sets & LLM-as-judge scoring — red-teaming, prompt-injection testing, OWASP LLM Top 10.

LLM-AS-JUDGE · RED-TEAM · OWASP
LLM EVALUATION    FINE-TUNING    AGENTIC SYSTEMS    AI SECURITY    RETRIEVAL & MEMORY    
( 03 )HOW THE WORK GETS BUILT

From raw data to deployed system.

01 — INGEST

Ingest & Ground

Parse corpora, OCR & structure into DB/CSV-grounded sources.

02 — TRAIN

Train & Fine-Tune

LoRA / QLoRA adaptation, DeepSpeed ZeRO runs & dataset curation.

03 — ORCHESTRATE

Orchestrate

Wire agents, tools & routing into a reliable workflow.

04 — SHIP

Ship & Harden

Eval harnesses, red-teaming & regression gates before every deploy.

( 04 )MOST RECENT WORK

Public Health Assistant

LIVE ON AWS
@ ARIZONA STATE UNIVERSITY · JUN 2024 — PRESENT · RESEARCH COLLABORATION

An evidence-grounded health assistant built with Dr. Mohan Tanirru (ASU). I fine-tuned open-source LLMs with LoRA/QLoRA and used frontier models to judge on clinical-dialogue data, own the evaluation pipeline — groundedness scoring, hallucination detection, regression gates — and red-team it against OWASP LLM Top 10, with automatic clinician-referral escalation on high-risk queries.

LORA / QLORA FINE-TUNESGROUNDEDNESS EVALS & REGRESSION GATESRED-TEAMING · OWASP LLM TOP 10RETRIEVAL & MEMORY ARCHITECTUREQUANTIZED GPU INFERENCE · AWS
FULL EXPERIENCE →
( 05 ) — SELECTED WORK

A few favourites.

ALL PROJECTS →
01 / AGENTS

AutoSteer

Multi-agent orchestration routing each request to the right specialist — 42 agents, config-driven.

PYTHON · MULTI-AGENT · YAML
02 / RAG

ChemRAG

Compliance agent with hybrid RAG, grounding guardrails and LLM-as-judge evaluation.

RAG · LITELLM · LANGFUSE
03 / TOOLING

claude-multimodel

Swap LLM providers inside Claude Code — a published npm package, global install.

TYPESCRIPT · NPM · CLI
( 06 )LET'S BUILD

Let's
build

GET IN TOUCH ↗