Site Reliability Engineer (SRE), Observability, London
Software Engineering
London, UK
- Deploy, support, and monitor new and existing services, platforms, and application stacks across production and non-production environments.
- Engage with our development and product partners to understand requirements, then design and implement resilient, scalable infrastructure solutions.
- Enhance, architect, author, and deliver software that improves the availability, scalability, and security of Apple's observability services.
- Build and run systems, infrastructure, and applications through automation â provisioning, configuration, deployment, and monitoring.
- Use scale testing to measure, tune, and optimise system performance; contribute to capacity planning and disaster-recovery exercises.
- Evaluate and integrate new technologies to improve system reliability, security, and performance.
- Collaborate on code, infrastructure, and design reviews, and drive process improvements.
- Participate in an on-call rotation, providing hands-on technical expertise during service-impacting events.
- Strong sense of ownership and integrity demonstrated through clear communication and collaboration
- Experience in managing and scaling distributed systems in a public, private, or hybrid cloud environment
- The ability to design, author, and release code in languages like (but not limited to) Go or Python
- Acute drive to automate manual operations and to improve them through repeated iteration
- Understanding of the Linux Operating System, standard networking protocols, and components
- Hands-on experience managing large numbers of diverse systems with configuration management or software delivery platforms (such as Puppet and Spinnaker)
- Experience with deploying, supporting and monitoring new and existing services, platforms, and application stacks
- Experience with scale testing, disaster recovery, and capacity planning
- Familiarity with microservices architecture and container orchestration with Kubernetes


