TopGenAIJobs

TopGenAIJobs

A Gen AI & Agentic AI Jobs Platform to discover high-quality Gen AI and Agentic AI opportunities from top companies worldwide.

topgenaijobs.com

Quick Links

  • Home
  • Browse Jobs
  • Browse by Category
  • Companies
  • Post a Job
  • Career Resources
  • About Us

Resources

  • Blog
  • Career Guide
  • Resume Tips
  • Interview Prep
  • Salary Guide
  • Skill Demand Index

Top Gen AI Roles

  • Gen AI Engineer Jobs
  • Agentic AI Engineer Jobs
  • Prompt Engineer Jobs
  • LLM Engineer Jobs
  • RAG Engineer Jobs
  • MLOps Engineer Jobs
  • Remote AI Jobs
  • Entry Level AI Jobs
  • Senior AI Jobs

Legal

  • Privacy Policy
  • Terms of Service
  • Cookie Policy
  • Contact

© 2026 TopGenAIJobs (Gen AI & Agentic AI Jobs Platform). All rights reserved.

Made with ❤ by TopGenAIJobs Team

    Home/Jobs/AI Platform Engineer

    AI Platform Engineer

    eBay

    Bengalore
    5+ years
    Today
    ₹23–44 LPA
    Full-time
    Hybrid

    Skills Required

    LLM Inference
    vLLM
    Python
    Go
    Rust
    PyTorch
    SGLang
    TensorRT
    GPU architecture
    CUDA
    Kernel Programming
    MLOps
    HPC
    Performance Optimization
    Quantization

    Description

    eBay is hiring an AI Platform Engineer focused on LLM inference for production systems. The role centers on making frontier-model serving fast, efficient, reliable, and observable.

    Company: eBay

    Role: AI Platform Engineer

    Location: Bengaluru, India | Hybrid

    Experience:

    • 5+ years of strong development experience
    • Experience deploying and operating LLM inference services in production
    • 3+ years hands-on experience in performance optimization and systems programming for AI/ML workloads
    • Strong production coding skills in Python plus Go or Rust
    • Demonstrated ability to deliver measurable production improvements
    • Proven skill in root-cause analysis across model, runtime, networking, and infrastructure

    Responsibilities:

    • Own production inference from handoff to production-grade serving
    • Handle release engineering, capacity planning, cost optimization, and incident response
    • Reduce end-to-end latency and increase throughput on real production traffic
    • Scale inference across heterogeneous GPU fleets
    • Build benchmarking suites, metrics, and tooling for latency, throughput, GPU utilization, memory, and cost
    • Improve monitoring, tracing, and alerting
    • Participate in incident response and postmortems
    • Evaluate research and implement pragmatic inference optimizations
    • Work with data science and product teams to translate business requirements into SLOs

    Additional responsibilities:

    • Optimize inference stacks and related components such as schedulers, KV cache, batching, and memory
    • Harden systems through operational learning and postmortem follow-up

    Nice to have:

    • CUDA/kernel programming
    • Speculative decoding

    More skills:

    GPU systems, continuous batching, KV cache management, speculative decoding, profiling, memory bandwidth, latency tradeoffs, monitoring, tracing, alerting, benchmarking, systems programming

    Other:

    • Engineering team
    • Equal opportunity employer
    • All qualified applicants receive consideration without regard to protected characteristics
    • Accommodation support is available on request
    • Talent Privacy Notice, Privacy Center, and AI Hiring Guidelines are referenced
    • AI tools may be used for administrative tasks in the hiring process
    • Accessibility statement is available

    Prepare for this role

    Recommended resources to build the skills for this position. Sponsored.

    Python for Everybody Specialization

    Coursera

    Learn Python from scratch — variables, data structures, web scraping, and databases.

    Python 3 Programming Specialization

    Coursera

    Intermediate Python covering classes, inheritance, APIs, and data processing.

    Deep Neural Networks with PyTorch

    Coursera

    Hands-on PyTorch from tensors to CNNs and transfer learning.