Infrastructure / Site Reliability Engineer (SRE)
Mercor (client confidential) · Remote
- Pay
- $200/hr
- Commitment
- part-time
- Hours / week
- ~40
- Source
- mercor
About this role
**Mercor** connects exceptional technical talent with leading organisations working on ambitious technology and AI initiatives. We are looking for experienced **Infrastructure / Site Reliability Engineers (SREs)** to join a full-time engagement focused on building and operating complex, enterprise-grade infrastructure. We are seeking engineers with strong hands-on experience building, operating, debugging, and scaling sophisticated production systems. The ideal candidate has worked extensively with **Kubernetes, AWS, observability platforms such as Datadog, and modern infrastructure tooling**. This is a **full-time opportunity**, and candidates must be able to commit to full-time engagement. ## What You'll Do - Build, operate, and improve highly available and scalable production infrastructure. - Manage and optimise **Kubernetes-based production environments**. - Design and maintain cloud infrastructure, primarily across **AWS**. - Improve system reliability, availability, scalability, and operational efficiency. - Build and maintain observability across infrastructure and applications using **Datadog** or similar platforms. - Investigate production incidents, perform root-cause analysis, and implement durable fixes. - Improve monitoring, alerting, logging, tracing, and overall production visibility. - Develop automation and internal tooling to reduce manual operational work. - Partner closely with software engineering teams on deployments, infrastructure, and production reliability. - Contribute to infrastructure architecture and technical decisions for complex distributed systems. ## Ideal Background - Professional experience in **Infrastructure Engineering, Site Reliability Engineering (SRE), Platform Engineering, DevOps, or Production Engineering**. - Hands-on experience operating **complex, enterprise-grade production systems**. - Strong production experience with **Kubernetes**. - Strong experience with **AWS** and cloud-native infrastructure. - Experience with **Datadog**, Prometheus, Grafana, or comparable observability platforms. - Experience with Infrastructure as Code using **Terraform, Pulumi, or equivalent technologies**. - Strong understanding of distributed systems, networking, containers, Linux, and cloud architecture. - Experience building or maintaining CI/CD and production deployment infrastructure. - Strong debugging, troubleshooting, and incident-response capabilities. - Proficiency in at least one programming or scripting language, such as **Python, Go, or Bash**. ## Strong Signals - Experience operating Kubernetes and cloud infrastructure at significant production scale. - Experience supporting high-traffic or mission-critical applications. - Experience building infrastructure or platform tooling used by large engineering organisations. - Ownership of production reliability, on-call operations, incident response, or capacity planning. - Experience working within sophisticated, large-scale distributed systems. - Demonstrated improvements to **SLOs/SLIs, observability, deployment reliability, infrastructure performance, or operational efficiency**. ### Why Join - Solve challenging reliability, scalability, and performance problems across enterprise-grade production systems. - Work extensively with technologies such as **Kubernetes, AWS, Datadog, Terraform/Pulumi**, and modern cloud-native tooling. - Take meaningful ownership of production reliability, observability, infrastructure architecture, and operational improvements. - Competitive hourly compensation reflecting your experience and technical expertise. - Join a network of highly skilled engineers working on ambitious projects with leading technology and AI organisations.
Skills & domains
- ai-training
- rlhf
- sme
- annotation
- Software Engineering

