FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

AI ML Ops Software Engineer
Smartsheet. Design, develop, and maintain stable and reliable AI/ML Ops platforms and pipelines .
Tech Stack
Tools & technologiesAWSAzureCloudDockerGoogle Cloud PlatformKubernetesPythonSQLTerraformUnity
About the role
Key responsibilities & impact- Design, develop, and maintain stable and reliable AI/ML Ops platforms and pipelines
- Package and deploy AI/ML services to production, ensuring reproducibility and interpretability
- Design and implement automated CI/CD pipelines to accelerate model deployment
- Provision and optimize infrastructure for model training and serving using Docker, Kubernetes, or serverless platforms
- Implement post-deployment monitoring for model performance, data drift, and latency
- Automate retraining and data pipeline workflows
- Manage deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation stacks, including vector databases and knowledge graphs
- Optimize GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference
- Collaborate with data scientists, data engineers, and software engineers to connect model development with production
- Manage versioning for data, code, and models using tools such as MLflow
- Implement data security measures and ensure compliance with data governance policies
- Evaluate emerging data technologies and identify infrastructure innovation opportunities
- Diagnose and resolve complex data-related issues
- Perform other duties as assigned
Requirements
What you’ll need- Minimum experience of 4–6 years in AI/ML Ops
- Experience building and maintaining scalable, reliable, efficient, and secure AI/ML Ops platform systems
- In-depth experience with AI/ML frameworks and tools involving large petabytes of data with the Databricks Lakehouse ecosystem
- Experience with AI/MLOps workflows on Databricks, MLflow, Mosaic AI Agent Framework, Unity Catalog, Vector Search, and Knowledge Graph
- Knowledge of AI/ML frameworks such as LangChain and LangGraph for AI/ML Ops pipeline integration
- Hands-on experience with at least one major cloud provider: AWS, Azure, or GCP
- Experience in an AWS-hosted data platform is preferable
- Proficiency in Python and SQL
- Experience with Kubernetes, CI/CD, infrastructure-as-code tools (preferably Terraform), observability, monitoring, and alerting
- Experience with enterprise SaaS software solutions requiring high availability and scalability
- Experience handling large-scale structured and unstructured data from varied data sources
- Experience with solution cost optimization and design-to-cost principles
- Legally eligible to work in India on an ongoing basis
- Experience with Monte Carlo is preferable
- Experience with AWS Bedrock is preferable
Benefits
Comp & perks- No specific benefits or compensation extras stated in the posting