Senior MLOps Engineer
Point Wild · Riga
Job description
About the role
As a Senior MLOps Engineer, you will play a critical role in architecting, building, and maintaining the infrastructure, pipelines, and tooling that enable complex AI models to be deployed, scaled, and monitored in production on Google Cloud Platform (GCP). You will collaborate closely with AI Researchers, Data Engineers, and Backend teams to bridge the gap between experimentation and high‑performance, enterprise‑grade production systems.
Key responsibilities
- Architect and manage scalable GCP‑based ML infrastructure using Vertex AI, GKE, Cloud Run, GCS and GPU/TPU instances.
- Own end‑to‑end model deployment lifecycle, building high‑throughput, low‑latency inference services with Docker, Kubernetes and serving frameworks such as Triton, vLLM or MLflow.
- Design automated CI/CD/CT pipelines for training, testing, evaluation and deployment using Airflow, Vertex AI Pipelines and GitHub Actions.
- Implement production observability and monitoring for system health and ML‑specific metrics (feature drift, prediction accuracy, data distribution shifts).
- Provide scalable training environments and standardized deployment templates for AI and research engineers.
- Collaborate with Data Engineers to integrate pipelines with feature stores, dataset versioning and batch/stream processing.
- Lead the transition of prototypes and notebooks into resilient, secure, auto‑scaling micro‑services.
Required profile
- At least 5 years of hands‑on experience designing, deploying and maintaining production ML workloads in cloud environments.
- Deep practical experience with GCP services including Vertex AI, Cloud Storage, GKE, Cloud Run and IAM/VPC.
- Expertise in containerization (Docker, Kubernetes) and model serving tools (Triton, vLLM, MLflow).
- Proven track record with workflow orchestrators (Airflow, Vertex AI Pipelines) and modern CI/CD tools (GitHub Actions, ArgoCD).
- Solid experience managing cloud resources as code using Terraform.
- Strong Python and SQL programming skills for scripting, automation and API development.
- Hands‑on experience with ML observability tools such as Grafana, Prometheus, GCP Cloud Monitoring or similar frameworks.
Required skills
- Google Cloud Platform (Vertex AI, GKE, Cloud Run, GCS)
- Docker
- Kubernetes
- Triton Inference Server
- vLLM
- MLflow
- Airflow
- Vertex AI Pipelines
- GitHub Actions
- ArgoCD
- Terraform
- Python
- SQL
- Grafana
- Prometheus
- GCP Cloud Monitoring
- Feature stores (Feast, Vertex AI Feature Store)
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches for Latvija.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published pirms 1 dienas
Expires pēc 1 mēneša
10 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Point Wild
Riga