Akamai Technologies Logo

Akamai Technologies

Senior Site Reliability Engineer

Reposted Yesterday
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Senior level
In-Office or Remote
Hiring Remotely in India
Senior level
Lead reliability and performance investigations across Akamai's global edge and media/web delivery platforms. Troubleshoot distributed systems across application, network, platform, and OS layers; design observability (SLIs/SLOs, telemetry, dashboards, alerts); analyze performance and bottlenecks; develop automation, internal tools, and AI-assisted diagnostics; partner with engineering and operations on scalable fixes; and provide escalation and off-hours support during critical incidents.
The summary above was generated by AI

Do you like collaborating across teams to solve complex problems?

Do you enjoy solving large scale distributed content delivery challenges?

Join our critical Edge Reliability Engineering Team!

Site Reliability Engineers at Akamai leverage software engineering, systems expertise, and operational skills to deliver reliable global services. The ERE team ensures performance, resilience, and availability of Akamai's media and web delivery platform while addressing distributed systems challenges. This role acts as the top technical escalation point for critical customer-impacting issues, connecting engineering, operations, and global support teams effectively.

Partner with the best

In this role, you will balance deep system diagnostics with progressive automation. You will serve as a premier technical authority for our platform and a champion for reducing systemic operational toil.

As a Senior Site Reliability Engineer, you will be responsible for:

  • Leading complex reliability and performance investigations across Akamai's global edge, media delivery, and web delivery platforms.
  • Troubleshooting critical distributed systems issues spanning application, platform, network, and operating system layers, serving as the highest technical escalation point.
  • Partnering with Engineering, Product, Support, and Network teams to identify root causes and deliver scalable, long-term solutions that improve platform reliability.
  • Designing and improving observability through SLIs, SLOs, KPIs, telemetry, dashboards, and alerts to identify and address customer-impacting issues.
  • Analyzing platform performance, traffic patterns, and system bottlenecks to improve scalability, resilience, and overall service reliability.
  • Developing automation, internal tools, AI-assisted diagnostics, and self-service workflows to streamline operations, reduce manual effort, and accelerate incident response.
  • Enhancing operational excellence through reliability-centered architecture reviews, post-incident analysis, continuous improvements, and offering off-hours support during critical incidents as needed.

Do what you love

To be successful in this role you will:

  • Possess Bachelors in CS/Engineering or a related field with 6 years of industry experience in large-scale SRE/Systems Infrastructure roles.
  • Have logical reasoning skills diagnosing complex performance bottlenecks, data integrity anomalies, and system failure modes in distributed environments.
  • Have understanding of internet technologies and foundational networking concepts, including caching, proxies, TLS, TCP/IP, DNS, and HTTP/HTTPS architectures.
  • Have foundation in Linux/Unix administration, diagnostic tools, and low-level environment troubleshooting.
  • Be able to retrieve data, analyze telemetry streams, and troubleshoot platform data integrity issues through SQL queries.
  • Have experience developing automation tools using languages like Python, Bash, or Go.
  • Demonstrate expertise in AI models and focus on implementing agentic workflows to reduce operational inefficiencies effectively.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Akamai Technologies Gurugram, Haryana, IND Office

Gurugram, India

Similar Jobs

7 Minutes Ago
In-Office or Remote
India
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
Oversee, scale, and optimize high-density AI hardware infrastructure across regional data centers. Build Python automation and infrastructure-as-code tooling, integrate incident workflows, develop telemetry pipelines and monitoring dashboards, and improve reliability across private cloud, bare-metal, and virtualized environments. Lead on-call incident response, runbooks, post-mortems, service rollouts, vendor coordination, and field technician activities while driving uptime, performance, and operational readiness.
Top Skills: Ai-Based Anomaly DetectionApi IntegrationsBare-Metal InfrastructureBgpGrafanaInfrastructure As CodeIpv4Ipv6LlmsLokiOpentelemetryPagerdutyPrivate CloudPrometheusPythonSlackTelemetry PipelinesVirtualization
7 Days Ago
Remote
Shri Bhrigukshetra, BLR, Uttar Pradesh, IND
Senior level
Senior level
Fintech • Analytics
Senior SRE responsible for service availability, performance, and scalability. Build automation and IaC, improve observability and reliability, participate in on-call rotations, incident response, postmortems, cloud migration enablement, and partner with development teams to improve release velocity.
Top Skills: AWSAzureBigpandaCi/CdDatadogDockerDynatraceEntraidGitKubernetesPythonShellTerraform
14 Days Ago
Remote
India
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software • Cybersecurity
Design, build, operate, and scale cloud-native infrastructure and Kubernetes clusters across major clouds. Implement CI/CD, observability, IaC, SLIs/SLOs, automate with Go/Python, participate in 24x7 on-call, and contribute to open-source and technical knowledge sharing.
Top Skills: AiopsAmazon EksArgo CdAzure AksCloudnativepgGithub ActionsGitlab Ci/CdGoGoogle GkeGpuGrafanaIstioJenkinsKubernetesLokiMimirNode.jsOpenshiftOpentelemetryPrometheusPulumiPythonTerraformThanos

What you need to know about the Delhi Tech Scene

Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account