ML DevOps Engineer
You have the unique opportunity to join our team, building an AI/ML platform to develop and train state-of-the-art AI-based cybersecurity agents. This is a hands-on infrastructure work in a security-sensitive environment, in collaboration with real external R&D teams building solutions on this very platform.
We are looking for an experienced engineer to help operate and enhance what has been done already to foster the platform growing even further. You will have the opportunity to collaborate with other internal project teams, and contribute in various ways to other project goals while ensuring scalability, compliance, and excellence in infrastructure services.
Key responsibilities
-
Design, build, and operate AI/ML infrastructure: GPU cluster management (Incus, LXD fork), Kubernetes (Talos Linux), model registry, inference serving, and data pipelines
-
Build reproducible, auditable environments with CI/CD and configuration management (Ansible, GitOps)
-
Provide the platform as an internal and external service: user on-boarding, support, documentation and capacity planning
-
Deploy, benchmark, and tune open-weight LLMs and other models on our own hardware
-
Work with R&D technical leads and project managers to keep infrastructure aligned with project content, budget, and schedule
-
Apply security-by-design in a classified environment: hardening, isolation, patching, compliance
-
Mentor junior team members and, as the team grows, take on team lead responsibilities, if necessary/possible
-
This role contributes to the Research and Development project AIDA, building into EU defence capabilities
Ideal candidate must have
-
3+ years of DevOps or infrastructure engineering experience in production environments
-
Solid fundamentals in Linux administration, Git, CI/CD, configuration management (e.g. Ansible)
-
Hands-on experience with containers and Kubernetes in production
-
Scripting and automation in Python and Bash
-
Experience operating services for other teams like monitoring, incident handling, documentation
-
Strong communication skills, both spoken and written, including proficiency in English
-
Willingness to apply for security clearance
-
Readiness to work on-site in Tallinn, with the option for partial hybrid work if desired
Ideal candidate would have experience in or would be willing to learn
-
ML workloads: GPU drivers/CUDA, model serving (vLLM, Ollama, llama.cpp), experiment tracking (Mlflow)
-
Infrastructure as code, GitOps, and observability tooling (Terraform/OpenTofu, Argo CD/Flux, KubeFlow, Prometheus/Grafana)
-
Cybersecurity practices, methodologies and standards
-
Working in restricted network environments
You don't need to tick every box, but if you're a strong DevOps engineer who is curious about ML/AI and wants to see your work used in real cybersecurity domain then we are looking for you!
We offer
-
Meaningful work
-
An innovative and driven work environment in a modern office in the heart of Tallinn, with dedicated desks for every employee
-
Motivated, supportive and fun team
-
International cyber security community
-
Continuous learning and professional growth opportunities
-
35 days of annual leave, and additional 1 extra day off each quarter, altogether 39 days of paid vacation a year
-
Employer-provided health insurance (after the probation period)
-
Stebby sports compensation (after the probation period)
-
Salary maintained during reservist training exercises
-
Average salary (based on the previous 6 months) paid for the first 3 days of sick leave
-
Enjoyable team events every few months!