Akamai Technologies Logo

Akamai Technologies

Senior Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Senior level
In-Office or Remote
Hiring Remotely in India
Senior level
Design, deploy, and maintain Akamai’s Compute platform infrastructure and internal tools. Improve service availability, reliability, scalability, observability, and performance through automation, monitoring, and configuration management. Support incident response and on-call operations, troubleshoot customer-impacting issues, define reliability requirements, and mentor other SRE engineers while collaborating across teams.
The summary above was generated by AI

Do you like collaborating across teams to solve complex problems?

Do you have a passion for cutting edge technologies and tackling system problems?

Join our highly-skilled Site Reliability team!

Our team designs, develops, and manages applications and infrastructure that support Akamai's Compute products and services. We create solutions that manage our Compute platform, focusing on cloud interfaces - Compute Portals and APIs. We do this while maintaining Akamai's mission to make life better for billions of people, billions of times a day.

Partner with the best

In this role, you'll ensure the operation and uptime of our Compute services and infrastructure. You'll supervise and maintain our critical infrastructure. You'll collaborate with cross-functional teams to create tooling and software that monitors and improves the reliability of our systems. You'll work with various technologies as we release brand new applications and modernize our existing tooling.

As a Senior Site Reliability Engineer, you will be responsible for:

  • Providing support and mentorship for other SRE engineers within the team
  • Defining requirements as part of the product lifecycle to influence the new designs and standards
  • Deploying and maintaining the platform and tools used internally
  • Partnering with multiple teams to ensure the availability, reliability, scalability and usability of our products and services
  • Improving our Compute Cloud Interface platform to speed error detection and remediation, enhancing performance and reliability
  • Developing and improving automation to support daily activities and reduce toil
  • Participating in on-call rotations, guiding restoration and repair of service-impacting issues
  • Collaborating with internal teams to help troubleshoot and resolve escalations and incidents for our customers.

Do what you love

To be successful in this role you will:

  • Have 5 years of relevant experience and a Bachelor's degree in Computer Science or its equivalent
  • Have experience automating with programming languages such as Python and/or Golang as well as scripting languages i.e. bash
  • Possess an understanding of best practices related to systems reliability, observability and monitoring incl. adherence to SLOs.
  • Have experience with configuration management tools such as SaltStack, Terraform and Ansible, as well as CI/CD solutions such as Jenkins
  • Have hands-on mastery in Linux administration and container-based platforms like Docker
  • Utilize monitoring tools like Prometheus, Grafana, Loki, web proxies (e.g., nginx/envoy/haproxy), and performance optimization solutions such as Redis effectively.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Akamai Technologies Gurugram, Haryana, IND Office

Gurugram, India

Similar Jobs

Yesterday
In-Office or Remote
India
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
The Senior Site Reliability Engineer improves the reliability, scalability, availability, and performance of distributed content delivery systems. Responsibilities include defining SLOs and SLIs, monitoring platforms, debugging incidents, implementing corrective actions, automating operational processes, participating in design reviews, and guiding scalable infrastructure design. The role collaborates with Product and Engineering teams and applies software engineering, systems administration, cloud, DevOps, and SRE practices.
Top Skills: AdbmsBashCloud ComputingDatadogDevOpsGrafanaJavaScriptOracle SqlPrometheusPythonUnix/Linux
10 Days Ago
In-Office or Remote
India
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
Oversee, scale, and optimize high-density AI hardware infrastructure across regional data centers. Build Python automation and infrastructure-as-code tooling, integrate incident workflows, develop telemetry pipelines and monitoring dashboards, and improve reliability across private cloud, bare-metal, and virtualized environments. Lead on-call incident response, runbooks, post-mortems, service rollouts, vendor coordination, and field technician activities while driving uptime, performance, and operational readiness.
Top Skills: Ai-Based Anomaly DetectionApi IntegrationsBare-Metal InfrastructureBgpGrafanaInfrastructure As CodeIpv4Ipv6LlmsLokiOpentelemetryPagerdutyPrivate CloudPrometheusPythonSlackTelemetry PipelinesVirtualization
17 Days Ago
Remote
Shri Bhrigukshetra, BLR, Uttar Pradesh, IND
Senior level
Senior level
Fintech • Analytics
Senior SRE responsible for service availability, performance, and scalability. Build automation and IaC, improve observability and reliability, participate in on-call rotations, incident response, postmortems, cloud migration enablement, and partner with development teams to improve release velocity.
Top Skills: AWSAzureBigpandaCi/CdDatadogDockerDynatraceEntraidGitKubernetesPythonShellTerraform

What you need to know about the Delhi Tech Scene

Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account