TopGenAIJobs

TopGenAIJobs

A Gen AI & Agentic AI Jobs Platform to discover high-quality Gen AI and Agentic AI opportunities from top companies worldwide.

topgenaijobs.com

Quick Links

  • Home
  • Browse Jobs
  • Browse by Category
  • Companies
  • Post a Job
  • Career Resources
  • About Us

Resources

  • Blog
  • Career Guide
  • Resume Tips
  • Interview Prep
  • Salary Guide
  • Skill Demand Index

Top Gen AI Roles

  • Gen AI Engineer Jobs
  • Agentic AI Engineer Jobs
  • Prompt Engineer Jobs
  • LLM Engineer Jobs
  • RAG Engineer Jobs
  • MLOps Engineer Jobs
  • Remote AI Jobs
  • Entry Level AI Jobs
  • Senior AI Jobs

Legal

  • Privacy Policy
  • Terms of Service
  • Cookie Policy
  • Contact

© 2026 TopGenAIJobs (Gen AI & Agentic AI Jobs Platform). All rights reserved.

Made with ❤ by TopGenAIJobs Team

    Home/Jobs/Senior AI/ML Developer

    Senior AI/ML Developer

    Innodata India Private Limited

    Noida
    4-7 years
    1 day ago
    ₹15–28 LPA
    Full-time
    Onsite

    Skills Required

    LLM
    Supervised Fine-Tuning
    Mistral
    vLLM
    Transformers
    RAG
    RLAIF
    DPO
    PPO
    GRPO
    RLHF
    LLaMA
    Qwen
    Constitutional AI
    Red teaming

    Description

    Seeking a Senior AI/ML Developer to lead RLHF pipeline development and optimization for advanced AI models.

    Role: Senior AI/ML Developer

    Location: Sector 62, Noida, Noida, Uttar Pradesh, India

    Experience:

    • 4 - 7 Years

    Responsibilities:

    • Own and drive full RLHF pipeline including data collection, reward model training, and RL fine-tuning
    • Design and run Supervised Fine-Tuning pipelines on open-weight models
    • Build and train reward models capturing human preferences
    • Design human feedback collection pipelines including labeling rubrics and annotator calibration
    • Implement Constitutional AI and RLAIF techniques to reduce human annotation reliance
    • Red team models post-training to identify jailbreaks, regressions, unsafe outputs, and alignment failures
    • Design and maintain evaluation benchmarks for alignment, safety, and capability
    • Optimize inference pipelines and runtimes to serve aligned models efficiently at scale
    • Implement quantization strategies to deploy fine-tuned models on target hardware
    • Write and tune low-level C/C++ and Rust code for inference performance
    • Diagnose and resolve training instabilities, reward hacking, and production inference bugs
    • Stay updated with latest alignment and RL research and translate findings into experiments

    Nice to have:

    • Experience with Direct Preference Optimization (DPO) and its variants
    • Understanding of reward hacking and Goodhart’s Law mitigation strategies
    • Experience diagnosing training instabilities such as reward collapse and mode collapse
    • Strong mathematical foundation in RL theory, probability, linear algebra, and optimization
    • Familiarity with transformer attention and tokenization
    • Experience with distributed training for large-scale fine-tuning
    • Experience with vector databases like FAISS or Milvus
    • Familiarity with retrieval-augmented generation (RAG) pipelines
    • Experience integrating LLMs with external tools, APIs, and agent-based systems
    • Exposure to Rapid Application Development (RAD) approaches

    More skills:

    Reinforcement Learning from Human Feedback (RLHF), Policy gradient methods, Supervised Fine-Tuning (SFT), Reward model training, Human feedback collection, Evaluation benchmarks, Inference optimization, llama.cpp, TensorRT, Quantization (INT4, INT8, FP8, LoRA, QLoRA), C, C++, Rust, Python, Distributed training frameworks, Vector databases (FAISS, Milvus), Retrieval-augmented generation (RAG), LLM integration with external tools and APIs, Rapid Application Development (RAD)

    Other:

    • Probationary employment

    Prepare for this role

    Recommended resources to build the skills for this position. Sponsored.

    Generative AI with Large Language Models

    Coursera

    Comprehensive LLM course covering transformer architecture, fine-tuning, RLHF, and deployment.

    Large Language Models: Application through Production

    edX

    Production-focused LLM course covering deployment, monitoring, and scaling.

    LangChain Chat with Your Data

    Coursera

    Build RAG applications with LangChain — document loading, splitting, embeddings, and retrieval.

    More LLM jobs

    Tech Lead AI Engineer

    Southwest Airlines

    Hyderabad

    Today

    Lead - Applied AI

    ORF

    Delhi

    Today

    AI Builder: Models & Agents

    Codewave

    Bengaluru

    Today

    AI Engineer

    GW RhythmX

    Bengaluru

    Today

    AI Engineer II

    Arintra

    Bengaluru

    Today

    Senior Backend Engineer (Agentic AI/LLM)

    CAI Stack

    Bengaluru

    Today