kjgpta.github.io

Kshitij Gupta


Forward Deployed EngineerML ResearcherBengaluru

I ship production LLM systems, and publish the NLP research that sharpens how those systems are built.

Selected

Experience

Where the work shipped

June 2026 – Present

Forward Deployed Engineer

TrueFoundry, Bengaluru

  • Partner with customers to ship production AI apps on cloud-native LLM infrastructure.
  • Turn platform capabilities into adoption paths and measurable business outcomes, not slide-deck demos.

July 2023 – May 2026

Machine Learning Engineer

Chubb Engineering Center India, Hyderabad

  • LLaMA-3.1 70B with LoRA/QLoRA and RAG: +25% accuracy, −15% drift.
  • Multi-agent planner / retriever / verifier: +18% factual grounding.
  • vLLM on AKS (A100/H100): −40% p95, +50% throughput, 10K+ daily requests.
  • Built HawkHire, an explainable AI hiring copilot for internal recruiting.

June 2022 – June 2023

NLP Research Intern

Speech Lab, NTU Singapore

  • English–Malay code-switched language models: +20% over baselines.
  • Published multilingual and code-switching work across ACL-adjacent venues.

2019 – 2023

B.E. Electrical & Electronics

BITS Pilani, Pilani Campus

  • Foundations in systems, ML coursework, and applied software projects.

Full résumé (PDF)

Featured work

tracesage

Local-first tracing for LangChain and LangGraph agents: one callback, a live graph, no cloud account.

tracesage showing a live LangGraph agent topology with agents, MCP tools, and LLM nodes
Fig. 1 Live agent topology, drawn from a single LangChain callback: agents, MCP tools, and LLM nodes with the run timeline beneath.

Problem Hosted tracers work well, but reaching for a cloud account to debug an agent at your own desk is the wrong shape of tool.

Approach Capture the LangChain callback stream, store each run in SQLite, and render a live graph and timeline locally. Prompts never leave the machine. A crash-safe handler, MCP tool-source attribution, a pytest fixture, and optional OpenTelemetry export round it out.

Outcome pip install "tracesage[langchain]" then tracesage demo. MIT-licensed and in beta; picked up by Python Weekly #750.

Python · LangChain · LangGraph · MCP · OpenTelemetry · SQLite

Research

Publications

Multilingual NLP, code-switching, and evaluation benchmarks: methods that make LLM systems sharper in the wild.

  1. arXiv · December 2024

    [1] WhoDunit: Evaluation Benchmark for Culprit Detection in Mystery Stories. A long-form narrative reasoning benchmark testing whether models can name the culprit across a whole story, not a short QA snippet. arXiv:2502.07747

  2. AACL-IJCNLP 2022 · Taiwan

    [2] MALM: Mixing Augmented Language Modeling for Zero-Shot Machine Translation. Mixing-based augmentation that improves zero-shot translation without parallel data for every language pair. In NLP4DH.

  3. ACIIDS 2023 · Thailand

    [3] Adapting Code-Switching Language Models with Statistical-Based Text Augmentation. Statistical augmentation for adapting language models to code-switched speech and text where labelled data is scarce. In ACIIDS.

  4. IALP 2023 · Singapore

    [4] Singaporean Conversational English-Malay Code-Switching Points. Characterises where English–Malay switches actually occur in Singaporean conversation, giving useful priors for code-switching models. In IALP.

  5. AISC 2023 · India

    [5] Data Augmentation for Automated Essay Scoring using Transformer Models. Augmentation strategies that hold up when scored essays are in short supply. In AISC.

Citations on Google Scholar

Writing

Notes on agents & tracing

Longer-form pieces on tracesage and local-first LangGraph observability.

All posts on Substack

Along the way

Earlier projects

Coursework and side projects that led into production ML and research.

Skills

Stack I reach for

Grouped by how I actually use it: model work, serving, and research.

Models & agents
Python, PyTorch, Transformers, LoRA and QLoRA, RAG, multi-agent systems, LangChain and LangGraph, Hugging Face.
Serving & infra
vLLM, Kubernetes, AKS, Azure, AWS, Docker, OpenTelemetry, Databricks, CI/CD.
Research
Natural language processing, code-switching, evaluation benchmarks.

About

How I like to work

I care about systems that survive real traffic, and methods that hold up under scrutiny.

Ship
Production over theatre. Latency, accuracy, and grounding under load, not prototype applause.
Prove
Measure what matters. Quantify the impact so a team can trust the system rather than the pitch.
Publish
Sharpen the method. When a problem needs a better evaluation or a better model story, write it down.

Open to collaborations on LLM tooling, production AI systems, agent observability, and applied NLP. Best reached by email.


Kshitij Gupta
mailguptakshitij@gmail.com