Senior MLOps / ML Platform Engineer
- Full-time
Company Description
We are looking for a Senior MLOps Engineer to join Sigma Software and help build a production-grade ML platform for a large-scale AdTech ecosystem. You will work on infrastructure powering predictive decision-making systems that process hundreds of millions of auction requests daily.
As part of a dedicated engineering team, you will contribute to scalable ML orchestration, model lifecycle automation, observability, and real-time optimization workflows. This role is ideal for engineers with strong production experience who enjoy solving complex platform and operational challenges.
We at Sigma Software offer the opportunity to work on cutting-edge ML infrastructure projects, collaborate with experienced engineers, and influence architecture decisions in a long-term strategic engagement.
CUSTOMER
Our Customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem. The company manages a high-load ad exchange platform handling hundreds of millions of auction requests every day and is investing in advanced predictive decision-making capabilities to improve advertiser outcomes and real-time optimization processes.
PROJECT
Sigma Software is building a predictive modeling and optimization platform integrated with a live ad exchange environment. The solution enables real-time supply scoring and filtering, audience look-alike generation, contextual performance estimation, and multi-objective optimization under operational constraints.
The project combines large-scale ML infrastructure, automated model lifecycle management, multi-tenant architecture, and advanced observability practices. The team focuses on delivering reliable, reproducible, and scalable ML systems ready for long-term Customer ownership.
Key Technologies: Python, Kubernetes, Docker, GCP, Vertex AI, MLflow, Airflow, Kubeflow, Argo Workflows, Terraform
Job Description
- Build and maintain ML training orchestration pipelines across hourly, daily, and weekly schedules
- Implement retries, backfills, and idempotent execution mechanisms
- Design and support model registry workflows including versioning, lineage, evaluation gates, and promotion processes
- Develop isolated per-advertiser model environments with namespace and configuration separation
- Build scalable refresh pipelines and publishing workflows for serving infrastructure
- Implement shadow mode and champion/challenger deployment strategies
- Develop monitoring and alerting for ML-specific metrics including feature drift, prediction drift, train/serve skew, and calibration decay
- Ensure reproducibility of ML workflows using containerized environments, pinned dependencies, and data snapshots
- Monitor training and scoring costs across tenants
- Collaborate with DevOps and SRE engineers on CI/CD and infrastructure automation
- Prepare operational documentation and platform handover materials
Qualifications
- 5+ years of experience in MLOps, ML platform engineering, or infrastructure engineering supporting production ML systems
- Strong Python skills and experience building platform-level tooling and automation
- Hands-on experience with Kubernetes and Docker
- Experience building CI/CD pipelines for ML workloads
- Hands-on production experience with MLflow, Kubeflow, Airflow, Argo Workflows, Vertex Pipelines, or similar orchestration and ML lifecycle platforms
- Experience with ML platforms and model lifecycle tools such as Vertex AI, MLflow, or Kubeflow
- Strong understanding of ML observability including drift detection, train/serve skew monitoring, and incident response
- Experience designing or supporting multi-tenant ML systems and isolated model environments
- Experience working with cloud platforms, preferably GCP
- Experience with infrastructure-as-code tools such as Terraform
- Experience with Linux environments
- Understanding of the ML lifecycle and productionization processes
- Upper-Intermediate English level or higher
WILL BE A PLUS
- Experience with feature stores and feature consistency management
- Experience with large-scale batch scoring systems operating under freshness SLAs
- Familiarity with experiment tracking platforms and evaluation gates
- Experience with on-premises Kubernetes or bare-metal Linux infrastructure
- Knowledge of DVC, lakeFS, or other data versioning tools
- Experience with Bigtable, Redis, Aerospike, or similar low-latency serving databases
- GPU scheduling and training cost optimization experience
- Familiarity with SOC 2, ISO 27001, or GDPR-related compliance requirements
Additional Information
PERSONAL PROFILE
- Strong ownership mindset and focus on operational reliability
- Ability to work independently in complex distributed systems environments
- Strong collaboration and communication skills
- Analytical thinking with attention to scalability and maintainability
- Comfortable working in fast-paced product-oriented environments
By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply