ItJobs Logo
Home About us Conditions
vi en
Login Sign Up
Logo

IT Jobs

Close
  • Home
  • About us
  • Conditions
  • Privacy
  • Contact
  • eng vi
TOP JOBS
Creative Force
Backend Developer
Creative Force
Up to 80000000VND
Rakuten Fintech Vietnam
Mid/Sr DevOps Engineer
Rakuten Fintech Vietnam
Up to 3200USD
Viettel Post
Data Engineer
Viettel Post
Up to 3000USD
CodeHQ
Senior MS Dynamics 365 Developer
CodeHQ
Up to 3000USD
One Mount Group
IT Infrastructure Expert
One Mount Group
Up to 3000USD
One Mount Group
OT Security Engineer
One Mount Group
Up to 3000USD
Nakivo
QA Team Lead
Nakivo
Up to 3000USD
Crossian
Supply Chain Data Analyst
Crossian
Up to 2850USD
One Mount Group
Senior Database Administrator
One Mount Group
Up to 2500USD

Pizza Hut Digital & Technology

Waseco Building, 10 Pho Quang, TP Hồ Chí Minh

Company Size : 25-99

View more

Job Summary

  • 25-99
  • Product
  • Việt Nam

AI Support Engineer

Pizza Hut Digital & Technology

  • Tan Binh, TP Hồ Chí Minh
  • Negotiable
  •  Full Time
  •  English
  •  Experienced (Non-Manager)
1
1

  •  Posted:05/08/2026

  • Apply now
AI Support Engineer
Apply now
Technical Skill: Python , AI (Artificial Intelligence) , Machine Learning , MS SQL , Distributed Systems , Docker , Splunk , Observability , MS Azure , DevOps , Grafana , AWS , Root Cause Analysis , Cloudwatch , Kubernetes , Bash , DataDog , GCP , Terraform , AWS CloudFormation , Prometheus , Amazon Sagemaker , MLflow , Kubeflow , MLOps , DataOps , CI/CD , Vertex AI , Azure Machine Learning

Job description

Overview of job

About Yum!

Our story might surprise you. We are the world’s largest restaurant company, encompassing KFC, Taco Bell, and Habit Burger & Grill, but there is much more happening behind the scenes than frying chicken, baking pizzas, and serving tacos.

We connect customers with our brands through apps, websites, kiosks, point-of-sale systems, and other digital dining experiences. Behind those experiences is a growing ecosystem of data, artificial intelligence, and machine learning solutions that support restaurant operations and customer engagement around the world.

Job Summary

We are seeking an experienced AI Support Engineer with a Machine Learning Engineering focus to join our 24/7 AI operations team.

This role is responsible for supporting, maintaining, and improving production machine learning platforms, pipelines, APIs, and deployment environments. The engineer will independently investigate complex operational issues, coordinate incident response, and partner with Machine Learning Engineers, Data Scientists, AI Engineers, platform teams, and cloud engineering teams to restore service and improve system reliability.

In addition to responding to incidents, this role will help strengthen the operational maturity of our AI ecosystem by developing automation, improving observability, refining support processes, and identifying recurring issues that should be addressed through engineering changes.

The successful candidate will have hands-on experience supporting cloud-based production systems and will be comfortable operating across machine learning, infrastructure, software engineering, and DevOps domains.

Key Responsibilities

Production Operations and Reliability

  • Monitor and support production machine learning pipelines, inference services, APIs, feature-processing workflows, and deployment environments.
  • Independently diagnose and resolve moderately complex issues involving model inference, data dependencies, deployment failures, service availability, latency, capacity, and infrastructure performance.
  • Assess the operational impact and urgency of production issues and take appropriate action to restore services.
  • Perform detailed root cause analyses for incidents, document findings, and recommend corrective and preventive actions.
  • Identify recurring failure patterns and partner with engineering teams to implement permanent solutions.
  • Support production readiness reviews and validate that new AI and machine learning services meet operational support requirements before release.
  • Participate in an on-call rotation supporting a 24/7 production environment.

Incident and Problem Management

  • Serve as a primary responder for AI- and MLE-related incidents identified through monitoring platforms, automated alerts, or user reports.
  • Lead or coordinate the technical investigation of incidents within the role’s area of responsibility.
  • Engage appropriate engineering, data, infrastructure, and vendor teams when cross-functional support is required.
  • Maintain clear communication during incidents, including impact, status, mitigation actions, and expected next steps.
  • Ensure incidents are tracked through resolution and that follow-up actions are documented and completed.
  • Analyze operational metrics such as mean time to acknowledge, mean time to resolution, incident volume, availability, and recurring failure categories.
  • Provide recommendations to improve reliability, supportability, and incident response effectiveness.

Deployment and Platform Support

  • Support CI/CD workflows used to test, package, deploy, and promote machine learning models and AI services.
  • Troubleshoot issues involving containers, Kubernetes workloads, cloud resources, infrastructure-as-code deployments, configuration, networking, permissions, and service dependencies.
  • Assist with model and platform releases, including deployment validation, rollback support, and post-release monitoring.
  • Partner with Machine Learning Engineers to improve deployment patterns, environment consistency, scalability, and recoverability.
  • Help manage production configurations, operational dependencies, and service-level requirements across development, staging, and production environments.

Automation and Continuous Improvement

  • Design and implement scripts, utilities, and automated workflows that reduce manual support effort and improve response times.
  • Automate common operational activities such as service recovery, endpoint scaling, health validation, deployment checks, log collection, and resource-management tasks.
  • Improve monitoring, logging, dashboards, alerts, and operational telemetry to provide earlier identification of production issues.
  • Develop and maintain runbooks, troubleshooting guides, support procedures, and operational playbooks.
  • Recommend improvements to architecture, tooling, processes, and support models based on operational experience.
  • Contribute to reliability, resilience, capacity-planning, and disaster-recovery initiatives for AI and machine learning services.

Cross-Functional Collaboration

  • Partner with Machine Learning Engineers to support the deployment, maintenance, and optimization of production models and machine learning infrastructure.
  • Collaborate with Data Scientists and AI Engineers to ensure new solutions are supportable, observable, and production ready.
  • Work with cloud, DevOps, cybersecurity, data engineering, and enterprise support teams to resolve issues spanning multiple technology domains.
  • Provide technical guidance to junior support engineers and assist with knowledge transfer, troubleshooting practices, and operational procedures.
  • Communicate complex technical issues clearly to both technical and nontechnical stakeholders.

Attractive Benefits:

  • 100% salary during probation period
  • Annual Leave: 18 days/ year
  • 2 days WFH/ week
  • Five “Recharge Days” – Extra days, in addition to company holidays.
  • Flexible Friday afternoon
  • Full salary insurance
  • 13th-month bonus
  • 1 day off for birthday
  • Advanced health insurance (Generali)
  • Regular engagement activities: sport clubs, internal event…
  • Support Macbook and Monitor

Job Requirement

Qualifications

  • Bachelor's degree in computer science, Engineering, Information Technology, Data Science, or a related field, or equivalent practical experience.
  • Three or more years of experience in machine learning engineering, DevOps, site reliability engineering, cloud engineering, software production support, or AI system operations.
  • Experience independently supporting business-critical applications or services in a production environment.
  • Strong working knowledge of at least one major cloud platform, such as AWS, Google Cloud Platform, or Microsoft Azure.
  • Hands-on experience with container technologies and orchestration platforms such as Docker and Kubernetes.
  • Experience with machine learning platforms or MLOps tools such as MLflow, Kubeflow, Vertex AI, SageMaker, or Azure Machine Learning.
  • Proficiency with Python and SQL for troubleshooting, scripting, data investigation, and automation.
  • Experience supporting CI/CD pipelines and infrastructure-as-code tools such as Terraform, CloudFormation, or equivalent technologies.
  • Ability to investigate application logs, infrastructure metrics, distributed systems failures, data-quality issues, and service dependencies.
  • Experience with incident management, problem management, root cause analysis, and production change management practices.
  • Strong analytical, troubleshooting, and communication skills.
  • Ability to manage multiple priorities and make sound technical decisions in a fast-paced, 24/7 support environment.
  • Ability to participate in an on-call rotation, including occasional support outside standard business hours.

Preferred Qualifications

  • Experience with monitoring, observability, and alerting tools such as Prometheus, Grafana, Datadog, Splunk, Cloud Monitoring, or CloudWatch.
  • Experience developing production-quality automation using Python, Bash, or another scripting language.
  • Understanding the complete machine learning lifecycle, including data preparation, model training, validation, deployment, monitoring, retraining, and retirement.
  • Experience supporting real-time inference services, batch-scoring pipelines, generative AI applications, or AI agents.
  • Familiarity with service-level indicators, service-level objectives, error budgets, and site reliability engineering practices.
  • Experience with model monitoring, data drift, model drift, performance degradation, and AI-specific operational risks.
  • Familiarity with secure software development practices, identity and access management, secrets management, and cloud security controls.
  • Experience mentoring junior engineers or leading technical incident investigations.
  • Experience supporting globally distributed systems, teams, or customers.

Languages

    • English

    • Speaking: Intermediate - Reading: Intermediate - Writing: Intermediate

Technical Skill

  • Python
  • AI (Artificial Intelligence)
  • Machine Learning
  • MS SQL
  • Distributed Systems
  • Docker
  • Splunk
  • Observability
  • MS Azure
  • DevOps
  • Grafana
  • AWS
  • Root Cause Analysis
  • Cloudwatch
  • Kubernetes
  • Bash
  • DataDog
  • GCP
  • Terraform
  • AWS CloudFormation
  • Prometheus
  • Amazon Sagemaker
  • MLflow
  • Kubeflow
  • MLOps
  • DataOps
  • CI/CD
  • Vertex AI
  • Azure Machine Learning

COMPETENCES

  • Reliable
  • Working Independently
  • Change Management
  • Communication Skills
  • Analytic Skills

Search for the right jobs

BUSINESS PROFILE

Pizza Hut Digital & Technology committed to building innovative digital solutions.

Why Pizza Hut Digital & Technology?  

We love pizza. We eat it a lot. There’s no doubt about that and we’re proud of it. But what makes us different is that it’s our people that drive the success of our business.

Alongside KFC, Taco Bell, and The Habit Burger Grill, we are part of the Yum! Family; the world’s largest restaurant company with nearly 43,000 restaurants in over 140 countries. At Pizza Hut, we are on a journey to build the most loved global brand and the fastest growing in every country; we have big plans over the next 5 years to achieve explosive growth in a competitive and ever-growing market.  We want to make it easy to get the best pizza and digital is at the core of this. We have recently created a startup: Pizza Hut Digital & Technology committed to building innovative digital solutions, providing our customers with an exceptional experience. From building a world-class platform and clean mobile experiences we are pushing the boundaries to ensure we are a source of innovation.

MORE JOBS FROM THIS EMPLOYER

  • 25-99
  • Product
  • Việt Nam

Senior Frontend Engineer

Pizza Hut Digital & Technology

  • Tan Binh, TP Hồ Chí Minh
  • Negotiable
  •  Full Time
  •  Experienced (Non-Manager)
1
Posted: 25/06/2026
Skills: React Native, Mobile App, TypeScript, Observability, Golang, OAUTH, Progress, Swift, Crashlytics, AWS, Firebase, Redux, ReactJS, Kotlin, SentryOne, JWT, DataDog, Figma, fastlane, Performance tuning, Detox, AWS Cognito, Cypress, CI/CD, Zustand, Copilot, React Query, MCP, Claude, Cursor, Monorepo, Turborepo, Java
  • 25-99
  • Product
  • Việt Nam

Software Engineer

Pizza Hut Digital & Technology

  • Tan Binh, TP Hồ Chí Minh
  • Negotiable
  •  Full Time
  •  Experienced (Non-Manager)
1
Posted: 08/06/2026
Skills: React Native, iOS, Android, AI (Artificial Intelligence), Swift, Java, Objective C, Kotlin, RESTful API, React Hooks

Search for the right jobs

footer_logo

WHO WE ARE

ITJobs is founded in 2014 in Vietnam and the primary goal is grow to one of the leading specialists in recruitment and selection of IT staff in Asia.

  • READ MORE

Jobs from Ho Chi Minh

  • Java jobs
  • C# jobs
  • Tester jobs
  • iOS jobs
  • ASP.NET jobs

Jobs from Hanoi

  • C++ jobs
  • Java jobs
  • Linux jobs
  • SQL jobs
  • .NET jobs

Information

  • About Us
  • Conditions
  • Privacy
  • Contact Us

ITJobs © Copyright 2013-2021