Kyivstar.Tech
DevOps(ML / LLM Infrastructure)
Remote Senior $4.7k–$12.8k/moest.
Summary
Kyivstar.Tech is seeking a DevOps Engineer to design, build, and operate infrastructure for their LLM platform. This DevOps-first role focuses on maintaining reliable, scalable, and efficient ML infrastructure across GPU/TPU/CPU environments, including data pipelines, training, and inference systems.
What you'll do
- Design, build, and operate scalable ML infrastructure on GCP (GKE) supporting experimentation and production workloads for LLMs and NLP systems
- Manage Kubernetes-based environments: deployment, scaling, upgrades, and reliability across GPU/TPU/CPU pools
- Build and maintain CI/CD pipelines (GitHub Actions, Jenkins) to automate testing, training, and deployment of ML services
- Implement infrastructure as code (Terraform, Ansible) to provision and manage cloud resources securely and cost-efficiently
- Ensure observability of ML systems through monitoring, logging, and alerting for infrastructure and production inference workloads
- Collaborate with ML Engineers and Data Engineers to design reliable training and inference pipelines
- Optimize resource utilization and cost of training and serving infrastructure
- Troubleshoot and resolve issues across the ML platform from data pipelines to distributed training and deployments
- Contribute to engineering best practices through code reviews and continuous improvement
Requirements
- Experience: 4+ years in DevOps, Platform Engineering, or ML Infrastructure roles with strong understanding of production systems and distributed workloads
- Cloud & Infrastructure: Hands-on experience with GCP; experience with other major cloud platforms is a plus; strong understanding of cloud-native architectures
- Kubernetes & Containers: Solid experience with Docker and Kubernetes (preferably GKE); familiarity with Helm and Kubernetes networking
- CI/CD & Automation: Experience building and maintaining CI/CD pipelines (GitHub Actions, Jenkins, or similar)
- Workflow Orchestration: Experience with Airflow or similar tools
- Infrastructure as Code: Strong experience with Terraform (preferred) or similar tools
- Programming: Strong hands-on experience with Bash and/or Python scripting languages
- Observability & Reliability: Experience with monitoring and logging systems (Prometheus, Grafana); understanding of reliability and debugging in distributed systems
- ML Infrastructure: Familiarity with ML lifecycle and experience supporting ML workloads in production environments
- Collaboration: Ability to work closely with ML Engineers and Data Engineers
Conditions
Work Arrangement: Office or remote - your choice
Benefits & Perks:
- Remote onboarding
- Performance bonuses
- Training and learning opportunities through company library, internal resources, and partner programs
- Health and life insurance
- Wellbeing program and corporate psychologist
- Reimbursement of expenses for Kyivstar mobile communication