Senior Site Reliability Engineer (Cloud and Networking)
Do you want to own the reliability of cloud load balancing infrastructure that serves thousands of customers at global scale?
Are you a senior technical leader who can drive solutions across distributed teams while mentoring the engineers around you?
Join our Cloud Networking SRE TeamThe Cloud Networking SRE team (CNETSRE) is part of Akamai's Infrastructure Engineering & Operations (IE&O) organization. We design, deploy, and manage the reliability of Akamai's core cloud networking products — including NodeBalancer, our production L4/L7 load balancer, and NLB (Network Load Balancer), our next-generation high-throughput L4 load balancing platform. These products are foundational to the Akamai Cloud Compute platform, serving customer workloads across dozens of global regions.
Partner with the bestAs a Senior Site Reliability Engineer on the NodeBalancer and NLB stack, you'll own the operational reliability of two generations of load balancing infrastructure: the production NodeBalancer fleet and the next-generation NLB platform with its distributed forwarding architecture. You'll build and maintain observability frameworks, lead incident response for complex multi-service failures, drive safe deployment practices across phased global rollouts, and mentor SRE II engineers on the team. The load balancing platform is actively evolving — future iterations are expected to move toward container-orchestrated deployments, so your Kubernetes expertise will be directly relevant as this stack grows. You'll work closely with NodeBalancer Engineering, the Product Delivery Team, and peer SRE functions to shape how the NB/NLB stack evolves operationally.
As a Senior Site Reliability Engineer, you will be responsible for:
- Owning the SRE lifecycle for NodeBalancer and Network Load Balancer — from design reviews and pre-rollout readiness assessments through production sign-off and ongoing reliability management
- Designing and implementing SLO/SLI frameworks that reflect true customer experience for L4 and L7 load balancing services, and driving action when error budgets are at risk
- Building and maintaining observability pipelines for NB/NLB infrastructure, including Prometheus metrics from load balancing components and system-level sources, and Grafana dashboards that enable rapid incident triage
- Leading technical incident response for complex NB/NLB failures — BGP/VIP issues, failover failures, data plane degradations, and configuration problems — acting as the technical commander and driving root cause analysis and preventive follow-through
- Developing and automating safe deployment workflows for phased NB/NLB releases, including bake period monitoring, feature flag management, and GO/NO-GO validation across global datacenter rollouts
- Reviewing design documents, product requirement Documents and producing actionable SRE input on operational risks, capacity implications, Day-2 concerns, and product strategy gaps
- Building automation and tooling using Python or Go that reduces operational toil and improves team-wide operational capability
- Mentoring SRE II engineers on the NB team, providing hands-on technical guidance, code/config reviews, and raising the bar for the team's SRE practice
- Participating in an on-call rotation for NB/NLB production systems, responding to incidents and driving resolution for customer-facing load balancing infrastructure
- Participate in a scheduled, daytime-only on-call rotation to spearhead technical incident response and resolve complex NB/NLB failures..
To be successful in this role you will:
- Have extensive experience in SRE, platform engineering, or infrastructure engineering, working with large-scale distributed systems
- Demonstrate deep expertise with Linux networking fundamentals — routing, BGP, nftables/iptables, ARP, VXLAN — and comfort diagnosing at the packet level using tcpdump, netstat, and similar tools
- Have hands-on experience with L4/L7 load balancing technologies — including proxy-based or kernel-level load balancers — covering configuration, health checking, high availability, and failure modes at scale
- Show a track record of defining SLO/SLI frameworks, building observability platforms from scratch, and running incident management processes at scale
- Demonstrate expertise in Kubernetes and containerization at scale — including workload scheduling, networking (CNI, Services, ingress), resource management, and operating stateful or network-intensive workloads in a cluster environment
- Build automation and tooling using Python or Go, with infrastructure-as-code experience (SaltStack, Ansible, or Terraform) and strong deployment safety instincts
- Demonstrate 4+ years in SRE or infrastructure engineering, with at least 2 years at cloud scale
Learn what makes Akamai a great place to work
Connect with us on social and see what life at Akamai is like!
We power and protect life online, by solving the toughest challenges, together.
At Akamai, we're curious, innovative, collaborative and tenacious. We celebrate diversity of thought and we hold an unwavering belief that we can make a meaningful difference. Our teams use their global perspectives to put customers at the forefront of everything they do, so if you are people-centric, you'll thrive here.
Working for you
At Akamai, we will provide you with opportunities to grow, flourish, and achieve great things. Our benefit options are designed to meet your individual needs for today and in the future. We provide benefits surrounding all aspects of your life:
- Your health
- Your finances
- Your family
- Your time at work
- Your time pursuing other endeavors
Our benefit plan options are designed to meet your individual needs and budget, both today and in the future.
About us
Akamai powers and protects life online. Leading companies worldwide choose Akamai to build, deliver, and secure their digital experiences helping billions of people live, work, and play every day. With the world's most distributed compute platform from cloud to edge we make it easy for customers to develop and run applications, while we keep experiences closer to users and threats farther away.
Join us
Are you seeking an opportunity to make a real difference in a company with a global reach and exciting services and clients? Come join us and grow with a team of people who will energize and inspire you!
#LI-Remote