Back to all roles

AI Platform

/

Senior leadership

Staff Infra Engineer — LLM Serving

Inference & Serving Systems

Lead Full-Time Remote (India preferred) / Guwahati Permanent

Compensation

INR 45 - 75 LPA + equity.

Engagement

Full-Time

Permanent role. Full-time commitment. Primarily on-site, with limited remote flexibility.

Scope of role

Set direction within your domain. Build and mentor a team. Own outcomes at the function level.

01 — The role

Why this role exists at EduRankAI

Own the inference fleet. Sub-100ms first-token latency at our scale, on infrastructure we control. Quantisation, batching, KV cache, scheduling, autoscaling — everything between "model checkpoint" and "API response". You will choose the inference stack, defend the choice, and own the cost-per-token reporting that finance uses to plan. When something goes wrong at 2am, you are who the on-call calls.

02 — The work

What you will own

  • 01 Own the inference stack end-to-end (vLLM / SGLang / TensorRT-LLM).
  • 02 Own the quantisation pipeline (GPTQ / AWQ / FP8) and the regression eval that says it shipped safely.
  • 03 Continuous batching and KV-cache management — including the unpopular hard cases.
  • 04 Multi-region serving and autoscaling.
  • 05 Cost-per-token observability used by finance.
  • 06 SLO ownership for every AI feature that ships.

03 — The expertise

What we look for

Production inference at >1000 RPSCUDA / Triton kernelsNetworking + RDMA basicsKubernetes + HelmStrong cost-vs-latency intuitionProfiler-fluent (nsys, py-spy, torch.profiler)

04 — The bar

Who thrives here

  • You have operated an inference fleet that served >1000 RPS in production for at least 12 months.
  • You can describe a specific quantisation regression you caught and how.
  • You have written at least one CUDA or Triton kernel that shipped to production.
  • You read kernel logs the way other people read newspapers.
  • You can defend a one-pager that argues for owning the inference stack instead of buying it.

05 — Hiring process

What to expect after you apply

  1. 01

    Application review

    Every application is read personally within five business days. We respond either way.

  2. 02

    Take-home or live exercise

    Role-specific. Time-boxed. Real problems we are actually working on, not invented puzzles.

  3. 03

    Conversations

    Deep technical and values conversations with the team you would join. No trick questions. No panel ambushes.

  4. 04

    Offer or honest no

    If yes: digital offer letter, signed in-portal, transparent terms. If no: written feedback if you want it.

Before you start

What we will collect. What it costs. What we will not do with it.

We will collect

  • Name, email, phone — Account + application updates. No marketing.
  • Resume / portfolio link — Human review of your work.
  • Date + place of birth — Identity verification only.
  • Your written responses — Selection rubric. Read by humans.
  • Government ID (later) — Anti-fraud at offer / interview stage. Not at signup.

We will never

  • Sell your data
  • Share with third-party recruiters
  • Use for advertising
  • Train models on it
  • Send marketing email

Our situation

EduRankAI is a small, independent organization building long-term capabilities in educational intelligence, advanced AI systems, and research infrastructure. We take no advertiser money, no donations with strings attached, and no investor pressure on hiring decisions. Applying is free, and every application is read by a human — recruitment, technical, academic and leadership teams. It buys us the right to be honest.

Full transparency policy Questions? Email us

Ready to apply?

We read every application personally. If you are the right person for this role — regardless of pedigree, background, or where you are based — you will hear back from us within five business days.

Lead Full-Time

Staff Infra Engineer — LLM Serving

Apply →