Mlops Lead Engineer And AI Architecture

January 20, 2026

No location

Full-time

RemotoJOB

Apply
Descripción

Job Description

Job Description
The candidate will be responsible for defining the architecture, building automated and scalable infrastructure, and establishing technical standards that drive artificial intelligence solutions for millions of users. Additionally, you will collaborate with researchers, data scientists, ML engineers, and cloud architects to translate research into products, ensuring automation, governance, and performance across all AI/ML initiatives.
Responsibilities:
Design and build MLOps infrastructure on AWS using Terraform, with a focus on security, scalability and profitability.
Build end-to-end CI/CD pipelines using GitLab and Jenkins for training, validation, deployments and rollbacks.
Contain ML models with Docker and deploy them to Kubernetes using Helm; Collaborate in the design and management of the Kubernetes platform.
Implement and manage the observability stack, including performance, reliability and security metrics.
Make key technical decisions, establishing patterns, tools and best practices for ML operations.
Collaborate with research and development teams to convert research into functional products.
Requirements:
5+ years in a senior DevOps, SRE, or MLOps role with a focus on production systems.
Extensive experience in architecting and managing Kubernetes clusters in a production environment.
Demonstrated proficiency in at least one infrastructure as code (IaC) tool, preferably Terraform.
Proficiency in a systems-level scripting language, such as Python or Go.
Experience creating and maintaining CI/CD pipelines for critical production services.
Direct experience implementing and managing specific ML models such as Agentic AI, NLU, ASR or TTS.
Experience with ML workflow orchestration tools such as Kubeflow or Apache Airflow.
Familiarity with ML experiment tracking and model registration tools, such as MLflow or SageMaker Model Registry.
Experience implementing models on specialized hardware, including GPU, Inferentia or Trainium.

Salary to receive
To agree