Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
CloudFactory

MLOps Support Engineer

CloudFactory

. Provide Tier 1 and Tier 2 operational support for AI/ML solutions .

Posted 9/24/2026full-timeKathmandu • NepalMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in providing operational support for AI/ML solutions, including troubleshooting, monitoring, and incident management. Proficient in SQL, Python, and cloud platforms, with a strong focus on model performance and bias prevention.

Highest-signal resume keywords
Operational Support for AI/ML SolutionsSQL ProficiencyPython ScriptingMonitoring and Observability ToolsMLOps Tooling

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
SQLPythonBashMLOpsGitKubernetesAI/ML SystemsModel Performance MonitoringIncident ManagementData Pipeline Health Monitoring
Soft Skills
Troubleshooting SkillsCollaborative MindsetAttention to Detail
Tools & Technologies
AWSGCPAzureGrafanaPowerBINew RelicMLflowDatabricks
Industry Keywords
DevOpsSREPlatform SupportModel DriftBias Prevention

Tech Stack

Tools & technologies
AWSAzureCloudGoogle Cloud PlatformGrafanaKubernetesPythonSQL

About the role

Key responsibilities & impact
  • Provide Tier 1 and Tier 2 operational support for AI/ML solutions
  • Identify failed jobs, degraded pipelines, and performance anomalies
  • Triage incidents, investigate issues, and coordinate escalation to Tier 3 Engineering
  • Participate in on-call rotas once established
  • Validate successful completion of pipelines and jobs
  • Monitor data pipeline health, model execution, and basic performance metrics
  • Identify operational issues before they impact customers
  • Respond to or alert customers about model outages or issues
  • Support incident management, rollback, and recovery activities
  • Use and maintain runbooks and operational documentation
  • Work with Engineering to improve supportability and observability
  • Contribute to knowledge sharing to reduce single points of failure
  • Work within defined SLAs and support processes
  • Build quarterly business reviews on ML model health
  • Evaluate champion/challenger models for promotion decisions
  • Monitor model drift and performance degradation
  • Validate that model updates and added data do not introduce bias
  • Provide MLOps coverage through assigned shift rotas and rotational on-call work

Requirements

What you’ll need
  • Experience in operations, DevOps, SRE, or platform support roles
  • Strong troubleshooting skills in production environments
  • Proficiency in SQL and scripting, including Python and Bash, for developing and automating ML workflows
  • Familiarity with cloud-hosted systems, including AWS, GCP, or Azure, for cloud-based ML services
  • Solid understanding of Git and version control in collaborative development environments
  • Comfortable working from runbooks and structured processes
  • Exposure to AI/ML systems in production
  • Familiarity with monitoring and observability tools such as Grafana, PowerBI, or New Relic
  • Knowledge of MLOps tooling and data platforms such as MLflow and Databricks
  • Experience supporting customer-facing platforms
  • Knowledge of containerization, including Kubernetes, is a plus
  • Experience with LLM prompt engineering and troubleshooting
  • Early career in MLOps or ML Engineering
  • Background in computer science, informatics, or related fields
  • Passion for machine learning and AI, with enthusiasm for optimizing and maintaining ML models in production environments
  • Collaborative mindset and willingness to contribute to model improvement, A/B testing, and iterative development
  • Attention to detail regarding model performance, bias prevention, and model behavior

Benefits

Comp & perks
  • Platform for professional growth, impact, and community
  • Opportunities to earn with purpose, learn every day, and serve a mission
  • Global community and collaboration across diverse cultures and perspectives