Jobiglo

No results.

Senior AI Compute Infrastructure Engineer

Kraken

Senior 🇬🇧 English
accelerator clusters device plugins workload isolation scheduling primitives orchestration quota management vLLM Triton Inference Server TensorRT model serving

Job description

About the role

Payward is building a dedicated AI Compute and Infrastructure team to power the next generation of model training, inference, evaluation, and experimentation for Kraken. As a Senior AI Compute Infrastructure Engineer you will design, operate, and optimise GPU and accelerator clusters that enable fast, reliable, and cost‑efficient AI workloads across the exchange.

Key responsibilities

  • Own and operate GPU and accelerator clusters, including drivers, runtimes, kernels, device plugins, node configuration and workload isolation.
  • Design infrastructure that allows teams to run models locally on GPUs, reducing reliance on external providers.
  • Build and improve scheduling, orchestration, placement, quota management and utilisation systems for heterogeneous accelerator environments.
  • Optimise inference pipelines for latency, throughput and cost using frameworks such as vLLM, Triton Inference Server or TensorRT.
  • Partner with ML engineers and researchers to remove bottlenecks in training, batch and online inference, deployment and debugging.
  • Develop observability for GPU utilisation, memory pressure, queue depth, token throughput, request latency and spend.
  • Drive reliability through incident response, alerting, runbooks and post‑incident improvements.

Required profile

  • Senior‑level engineer with a strong background in building production‑grade compute infrastructure.
  • Experience working closely with AI/ML researchers, platform engineers, security and product teams.
  • Proven ability to deliver fast, dependable and cost‑effective solutions at scale.

Required skills

  • GPU and accelerator cluster management
  • Linux drivers, runtimes and kernel configuration
  • Device plugins and workload isolation techniques
  • Scheduling, orchestration and quota management
  • vLLM, Triton Inference Server, TensorRT or equivalent serving stacks
  • Model serving and inference pipeline optimisation
  • Observability and monitoring of GPU utilisation and performance metrics

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Kraken.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches for Latvija.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published pirms 4 nedēļām

Expires pēc 1 mēneša

22 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Kraken