SULAV.K.SHRESTHA // BOOT000
POST /v1/infer202 ACCEPTED

SULAV KUMAR
SHRESTHA

AI SYSTEMS ENGINEER

Building intelligent systems at the intersection of LLMs, inference optimization & cloud infrastructure.

STAGE 01 — QUEUE /batch/join

BATCHING

THE REQUEST JOINS THE QUEUE. CONTEXT ATTACHED.

I work on building AI-driven systems for real-world workflows in a compliance-focused environment. My work involves developing data validation pipelines, document processing systems, and internal automation tools powered by machine learning.

Over time, I've become more interested in what happens beyond model usage — how systems actually behave under load. Recently, I've been focusing on ML systems and LLM inference, working with Go to build backend services and gateways that handle request flow, concurrency, timeouts, and system limits.

INTERESTS

  • LLM serving and inference systems
  • Backend systems under real-world constraints
  • Batching, scheduling, and request control
  • Building reliable and cost-aware AI infrastructure

CURRENTLY

  • Shipping LLM workflows on Azure
  • Optimizing inference for edge and cloud
  • Designing scalable microservice AI systems

STAGE 02 — ROUTE svc-mesh

THE CLUSTER

ROUTED POD TO POD ACROSS THE SERVICE MESH.

DISTRICT A — PRODUCTION // AKS · SERVICE BUS

InCorp

AI & DATA ENGINEER — INTERN

APR 2025 — PRESENT · ONSITE

  • Built PDF-to-Excel extraction pipelines with Azure Document Intelligence + Azure OpenAI — cutting processing time from 1–2 hours to under 5 minutes.
  • Designed a document compliance system across 6 microservices / AKS pods with Azure Service Bus routing — classification, verification and validation over hundreds of thousands of documents.
  • Built transaction risk workflows flagging duplicates, anomalies and compliance violations.
FastAPICosmosDBRedisMS Graph APISharePointAzure OpenAIAzure Document IntelligenceAKS

DISTRICT B — VISION // EDGE · TRACKING

Sprhava

JR. DATA SCIENTIST — GERMANY, REMOTE

JUL 2024 — NOV 2024

  • Enhanced YOLOv8 obstacle tracking, reducing false positives in production streams.
  • Replaced DeepSORT with ByteTrack — +30% FPS on the same hardware budget.
  • Optimized the pipeline for cloud deployment.
YOLOv8ByteTrackPythonOpenCV

STAGE 03 — EXEC gpu-pool

GPU RACKS

KERNELS DISPATCHED. HOVER A RACK TO LIGHT IT UP.

LANGUAGES

compilers & runtimes

  • Python
  • Go
  • SQL

AI / ML

vision & numerics

  • YOLOv8
  • OpenCV
  • MediaPipe
  • ONNX Runtime
  • NumPy
  • Pandas

LLM / NLP

serving & orchestration

  • LangChain
  • Hugging Face Transformers
  • Azure OpenAI
  • Microsoft Foundry
  • vLLM
  • Ollama
  • Prompt Engineering

INFRASTRUCTURE

state, transport & compute

  • Docker
  • FastAPI
  • PostgreSQL
  • CosmosDB
  • Redis
  • Azure Service Bus
  • AKS

STAGE 05 — RESPONSE STREAM

RESPONSE

200 OK — INFERENCE COMPLETE · CONNECTION KEPT ALIVE