Eastern Peak
Site Reliability Engineer / DevOps Engineer
Remote Senior $4.7k–$12.8k/moest.
Summary
Eastern Peak is seeking a Site Reliability Engineer / DevOps Engineer to scale and operate core infrastructure for a rapidly growing AI infrastructure platform. You will join an expanding infrastructure team focused on building reliable, scalable, and cost-efficient systems powering production workloads across distributed environments. This role is ideal for someone who enjoys solving complex operational challenges and improving system reliability in a fast-moving startup.
What you'll do
- Build, operate, and continuously improve infrastructure powering a distributed inference platform
- Own reliability, scalability, and operational excellence across AWS-based control planes and multi-provider GPU fleet
- Design and maintain networking between control planes, Kubernetes clusters, and geographically distributed GPU hosts
- Operate and enhance Kubernetes-based orchestration (primarily EKS)
- Manage deployments and infrastructure using Helm, FluxCD, and Terraform
- Improve observability through metrics, logs, traces, dashboards, and alerting (Prometheus, Grafana, Loki, Jaeger, OpenTelemetry)
- Tune alerts, develop runbooks, and strengthen operational readiness as systems scale
- Respond to production incidents, perform root cause analysis, and implement long-term fixes
- Collaborate with globally distributed engineers using clear asynchronous communication and structured handoffs
Requirements
Required:
- 5+ years of experience in SRE, DevOps, platform engineering, or infrastructure engineering
- Strong production experience with Kubernetes and networking
- Hands-on experience operating AWS infrastructure, especially EKS
- Strong Linux and distributed systems experience, including environments beyond fully managed cloud abstractions
- Experience with Prometheus, Grafana, Loki, Jaeger, and OpenTelemetry
- Experience with GitOps and deployment workflows (Helm, FluxCD or similar)
- Infrastructure as Code experience, ideally Terraform
- Experience with alert tuning, runbooks, and incident management practices
- Clear communicator able to work effectively in async, distributed teams
- Excellent English (spoken and written)
Nice to Have:
- Experience with AI inference, ML infrastructure, or high-performance distributed systems
- Experience managing GPU fleets, bare-metal infrastructure, or multi-provider compute environments
- Experience using AI tools to improve engineering productivity
Conditions
Remote position with globally distributed team. Async-first work environment with clear communication and structured handoffs.