RocketRoute, an APG company
Site Reliability Engineer
Remote Senior $4.7k–$12.8k/moest.
Summary
RocketRoute, an APG company, is seeking a Site Reliability Engineer to ensure reliability, scalability, and performance of critical infrastructure systems. This leadership role focuses on implementing automation, optimizing system monitoring, and enhancing security and operational excellence.
What you'll do
- Design, implement, and maintain organizational infrastructure and systems for optimal performance and reliability, including self-hosted systems, server maintenance, and system upgrades
- Lead complex incident resolution with minimal downtime, drive incident response processes, and conduct root cause analysis and post-mortem reporting
- Implement and optimize system monitoring and alerting across infrastructure; proactively identify and resolve issues before affecting end-users
- Drive automation initiatives for server provisioning, configuration management, and incident response; build and enhance internal automation tools
- Collaborate with security teams to enforce best practices, ensure secure access management, and automate security patching and vulnerability management
- Work with development teams and DevOps to implement reliable systems and ensure smooth application deployments
- Oversee user account provisioning, de-provisioning, and SaaS platform management in coordination with IT and HR
- Design, implement, test, and continuously improve disaster recovery and backup processes
- Create and maintain comprehensive documentation; share best practices through knowledge-sharing sessions
- Evaluate and implement new technologies to improve system performance and reduce operational costs
Requirements
- 3+ years of experience in Site Reliability Engineering, DevOps, or IT operations with hands-on large-scale systems management
- Proficient with Linux/Unix systems, particularly RHEL and Ubuntu distributions
- Experience with AWS cloud infrastructure and ability to manage cloud-based environments
- Strong understanding of networking concepts (TCP/IP, DNS, HTTP/HTTPS) and troubleshooting experience
- Hands-on experience with Apache, Nginx, MySQL, PostgreSQL, Docker, Kubernetes, Zabbix, Ansible, Puppet, and Terraform
- Proficiency with infrastructure automation tools and configuration management best practices
- Strong scripting and automation skills in Python, Bash, or Ruby
- Experience with monitoring tools such as Prometheus, Grafana, or Datadog and incident management platforms
- Strong troubleshooting skills for complex incidents and system failures
- Good understanding of security best practices, access management, vulnerability scanning, and system hardening
- Experience with disaster recovery, backup, and high availability solutions
Conditions
Schedule: Full-time, Monday to Friday
Location: Argyle, TX - Onsite