TopGenAIJobs

TopGenAIJobs

A Gen AI & Agentic AI Jobs Platform to discover high-quality Gen AI and Agentic AI opportunities from top companies worldwide.

topgenaijobs.com

Quick Links

  • Home
  • Browse Jobs
  • Browse by Category
  • Companies
  • Post a Job
  • Career Resources
  • About Us

Resources

  • Blog
  • Career Guide
  • Resume Tips
  • Interview Prep
  • Salary Guide
  • Skill Demand Index

Top Gen AI Roles

  • Gen AI Engineer Jobs
  • Agentic AI Engineer Jobs
  • Prompt Engineer Jobs
  • LLM Engineer Jobs
  • RAG Engineer Jobs
  • MLOps Engineer Jobs
  • Remote AI Jobs
  • Entry Level AI Jobs
  • Senior AI Jobs

Legal

  • Privacy Policy
  • Terms of Service
  • Cookie Policy
  • Contact

© 2026 TopGenAIJobs (Gen AI & Agentic AI Jobs Platform). All rights reserved.

Made with ❤ by TopGenAIJobs Team

    Home/Jobs/AI Solutions and Platforms Operations Engineer

    AI Solutions and Platforms Operations Engineer

    PepsiCo

    Hyderabad
    3-5 years
    Today
    ₹15–28 LPA
    Full-time
    Onsite

    Skills Required

    LLM
    RAG
    LangChain
    Agentic frameworks
    Embeddings
    Vector Database
    Python
    APIs
    Microservices
    Vector search
    Azure
    AWS
    GCP
    Kubernetes
    OpenTelemetry

    Description

    The AI Observability Engineer develops and operationalizes agentic AI solutions using orchestration frameworks and contributes to an AI Agent Operations Center for safe, reliable, and observable agent behavior at scale.

    Company: PepsiCo

    Role: AI Solutions and Platforms Operations Engineer

    Location: Hyderabad, India

    Experience:

    • 3–5+ years of software engineering experience
    • 1+ years building and observing AI/ML or GenAI applications preferred

    Qualification:

    • Bachelor’s in Computer Science, AI/ML, Data Science, or a related field

    Responsibilities:

    • Build agent runtime management capabilities including agent registry, versioning, deployment tracking, and run histories
    • Enable operational workflows such as incident triage, replay/debug runs, trace correlation, and root-cause analysis
    • Implement operational dashboards for agent health metrics like success rate, latency, tool failure rate, cost per run, and loop detection
    • Instrument agent flows end-to-end using OpenTelemetry or equivalent
    • Implement semantic conventions and tagging standards
    • Partner with SRE/observability teams for production-grade monitoring, alerting, and operational readiness
    • Collaborate with transformation teams and business stakeholders to tailor AI agents to specific domains
    • Work closely with AI platform teams to build scalable and cross-domain AI agents with end-to-end observability
    • Build and maintain CI/CD pipelines for agent services and operations center components
    • Automate onboarding for new agent use cases
    • Drive best practices for secure, scalable, and cost-effective agent deployments
    • Stay updated with AI and machine learning advancements and integrate them into AI agents
    • Conduct testing and validation to ensure reliability and accuracy of AI agents

    Nice to have:

    • Experience building internal developer platforms or operational consoles
    • Strong engineering discipline including testing, versioning, CI/CD, and automation
    • Operational mindset focused on reliability, debuggability, and incident response support
    • Problem-solving skills to translate business challenges into technical solutions
    • Collaboration skills for effective cross-functional teamwork
    • Agility to adapt to changing requirements and new technologies
    • Communication skills to explain complex technical concepts to non-technical stakeholders

    More skills:

    agentic frameworks (Crew.ai, LangChain, Semantic Kernel, AutoGen, or similar), microservices patterns, RAG patterns (embeddings, vector search, retrieval evaluation, chunking strategies), cloud environments (Azure, AWS, GCP), containerized deployments (Kubernetes, AKS, EKS), observability fundamentals (logs, metrics, traces), distributed tracing, telemetry pipelines, Azure AI Search, prompt/version management, evaluation frameworks, Responsible AI practices (data handling, safety guardrails, audit trails, redaction strategies), FinOps (token/GPU cost optimization, chargeback/showback reporting)

    Prepare for this role

    Recommended resources to build the skills for this position. Sponsored.

    Python for Everybody Specialization

    Coursera

    Learn Python from scratch — variables, data structures, web scraping, and databases.

    Python 3 Programming Specialization

    Coursera

    Intermediate Python covering classes, inheritance, APIs, and data processing.

    Generative AI with Large Language Models

    Coursera

    Comprehensive LLM course covering transformer architecture, fine-tuning, RLHF, and deployment.

    More LLM jobs

    Solutions Architect- Python & AI/ML

    FloTorch

    Hyderabad

    Today

    Gen AI-AI and Data Science Engineer III

    Deloitte

    Bengaluru

    Today

    Model and Intelligence Engineer

    Ericsson

    Noida

    Today

    Model and Intelligence Engineer

    Ericsson

    Noida

    Today

    ML Engineer

    SynthBee Inc.

    Jammu

    Today

    Senior AI Engineer

    Backbase

    Hyderabad

    Today