hireclover
All jobs

Eastern Peak

Site Reliability Engineer / DevOps Engineer

Remote Senior $4.7k–$12.8k/moest.

Summary

Eastern Peak is seeking a Site Reliability Engineer / DevOps Engineer to scale and operate core infrastructure for a rapidly growing AI infrastructure platform. You will join an expanding infrastructure team focused on building reliable, scalable, and cost-efficient systems powering production workloads across distributed environments. This role is ideal for someone who enjoys solving complex operational challenges and improving system reliability in a fast-moving startup.

What you'll do

  • Build, operate, and continuously improve infrastructure powering a distributed inference platform
  • Own reliability, scalability, and operational excellence across AWS-based control planes and multi-provider GPU fleet
  • Design and maintain networking between control planes, Kubernetes clusters, and geographically distributed GPU hosts
  • Operate and enhance Kubernetes-based orchestration (primarily EKS)
  • Manage deployments and infrastructure using Helm, FluxCD, and Terraform
  • Improve observability through metrics, logs, traces, dashboards, and alerting (Prometheus, Grafana, Loki, Jaeger, OpenTelemetry)
  • Tune alerts, develop runbooks, and strengthen operational readiness as systems scale
  • Respond to production incidents, perform root cause analysis, and implement long-term fixes
  • Collaborate with globally distributed engineers using clear asynchronous communication and structured handoffs

Requirements

Required:

  • 5+ years of experience in SRE, DevOps, platform engineering, or infrastructure engineering
  • Strong production experience with Kubernetes and networking
  • Hands-on experience operating AWS infrastructure, especially EKS
  • Strong Linux and distributed systems experience, including environments beyond fully managed cloud abstractions
  • Experience with Prometheus, Grafana, Loki, Jaeger, and OpenTelemetry
  • Experience with GitOps and deployment workflows (Helm, FluxCD or similar)
  • Infrastructure as Code experience, ideally Terraform
  • Experience with alert tuning, runbooks, and incident management practices
  • Clear communicator able to work effectively in async, distributed teams
  • Excellent English (spoken and written)

Nice to Have:

  • Experience with AI inference, ML infrastructure, or high-performance distributed systems
  • Experience managing GPU fleets, bare-metal infrastructure, or multi-provider compute environments
  • Experience using AI tools to improve engineering productivity

Conditions

Remote position with globally distributed team. Async-first work environment with clear communication and structured handoffs.

Browse jobs