Connect with me →
[ Pratham Kotkar ]-> backend & ai systems engineer

I build the backend behind AI products, and make it hold under load.

Shipped at

Ardent Privacy, enterprise data discovery

Built

An LLM gateway in Go that holds 5,000 req/s

Now

M.S. Computer Science, UC Santa Cruz

GoPythonNode.jsRust PostgreSQLRedisDockerAWSLangGraph
30%

faster API responses on the Data Discovery backend, from query tuning and Redis caching

2–10ms

p95 overhead added by my LLM gateway at 5,000 requests a second

64→83%

recall@5 on a RAG service over 7,600 documents, versus vector-only search

<140ms

to fail over to another model provider during simulated outages

What I’ve built, and what it holds up under

Selected work, 2024–2026
AP

Ardent Privacy

Data Discovery · software engineer

Backend for the flagship product that lets enterprises scan, classify and manage sensitive data, with compliance logic and the APIs the frontend runs on.

  • Modular REST APIs in Node.js and Express, backed by CockroachDB.
  • Query optimization and Redis caching, scaled to millions of records.
  • CI/CD pipelines, Docker containers and release deployments.
30% faster responsesMillions of records
Mar 2024 – Mar 2026
GW

LLM Gateway

Multi-provider routing · personal project

Apps that call model APIs directly inherit provider outages, runaway costs and leaked sensitive data. This gateway sits between services and providers and absorbs all three.

  • Provider routing with automatic failover and SSE streaming passthrough.
  • Redis token-bucket rate limiting and semantic response caching.
  • PII redaction middleware and Prometheus metrics for cost and latency.
+2–10 ms p95 at 5,000 req/s−32% repeat-query spend<140 ms failover
Go · Redis · Prometheus · Docker · Sep 2025 – Jan 2026
RAG

Production RAG Service

Retrieval with an evaluation harness

Keyword search misses context and LLM answers go unverified. This API answers over 7,600 documents and scores its own quality.

  • Async ingestion workers and hybrid vector plus full-text retrieval with reranking.
  • Redis response caching behind FastAPI.
  • Eval harness measuring recall@k and answer faithfulness on a labeled question set.
recall@5 64% → 83%167 ms p9559% cache hits
FastAPI · PostgreSQL/pgvector · Redis · Docker · Oct 2025 – Feb 2026
IR

AI Incident Response Platform

Multi-agent workflow

A distributed incident response platform that orchestrates several AI agents through a checkpointed workflow that can resume where it stopped.

  • Asynchronous coordination services in FastAPI.
  • Approval workflows with safety guardrails before agents act.
  • LangSmith observability, and LLM reasoning checked by heuristic validation.
Checkpointed agentsHuman approval gates
FastAPI · LangGraph · LangSmith · Mar – Apr 2026

Most LLM systems don’t fail at the model. They fail at the plumbing.

Approach

Routing, rate limits, caching, redaction, failover. These are the parts nobody demos, and the parts that decide whether a product stays up. My gateway handles them in one place so each service does not have to.

Backend engineer, moving deeper into AI infrastructure

About
Portrait of Pratham Kotkar

I spent two years at Ardent Privacy building the backend for a data discovery product, where sensitive data, compliance rules and large record counts leave little room for slow or fragile services.

Outside that job I build the infrastructure around language models: a gateway that survives provider outages, a retrieval service that measures its own answer quality, and an agent workflow with approval gates.

I’m now studying for an M.S. in Computer Science at UC Santa Cruz, after a B.E. in Computer Science from Savitribai Phule Pune University (3.81/4.00).

Experience

Software Engineer, Ardent PrivacyMAR 2024 – MAR 2026

Backend and APIs for the Data Discovery product, plus CI/CD and deployments.

Software Development Intern, MandrakeTechAUG – NOV 2023, PUNE

Built a mock Gmail API in Rust with the Gmail API and OAuth, and learned Docker deployment.

Software Development Intern, BlueBinariesFEB – MAY 2023, PUNE

Flutter and Firebase app that explains dashboard warning lights. Cut issue identification from 55 to 25 seconds.

Education

M.S. Computer Science, UC Santa CruzSEP 2026 – PRESENT
B.E. Computer Science, Savitribai Phule Pune UniversityAUG 2020 – MAY 2024 · 3.81/4.00

Algorithms, advanced data structures, network security, machine learning, NLP, data visualization.

Research & recognition

Smart Workout CompanionPAPER · MAY 2024

Deep learning for posture monitoring and yoga pose refinement. Presented as “PoseCraft Fitness Navigator” at ICRTACT-2024, an international conference.

OOP Hackathon runner-upMMCOE · DEC 2021

Skills

Languages
Go, Python, Rust, JavaScript, C, C++
Backend & data
Node.js, Express.js, Flask, CockroachDB, PostgreSQL, MySQL, MongoDB, Redis
Cloud & DevOps
AWS, Azure, GCP, Docker, Jenkins, Linux, Git, CI/CD
Vision
OpenCV, MediaPipe

Certifications & leadership

AWS Academy Cloud FoundationsAMAZON WEB SERVICES
IBM Cybersecurity FundamentalsIBM SKILLSBUILD
Core committee, Google Developer Student ClubSOFTWARE DEVELOPMENT, MMCOE
Software team member, TEAM CODEMMCOE
Contact

Tell me what you’re building.