Akamai Technologies Logo

Akamai Technologies

Senior Site Reliability Engineer

Reposted 2 Hours Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Senior level
In-Office or Remote
Hiring Remotely in India
Senior level
Senior SRE responsible for monitoring, analyzing, and improving availability, performance, and reliability of Akamai's Mapping Service. Define KPIs, build tooling to prevent recurrence, collaborate with product engineers on scalable designs, troubleshoot incidents, and use data analysis and network diagnostics to recommend improvements.
The summary above was generated by AI

Do you like collaborating across teams to solve complex problems?

Do you enjoy solving large scale distributed systems problems?

Join the Mapping SRE team

As Akamai Mapping SREs, we manage the reliability, performance, and scalability of a global system routing trillions of daily client requests and tens of terabits of traffic per second. We combine core SRE principles with data science to analyze massive datasets, define critical KPIs, build advanced monitoring infrastructure, and troubleshoot complex production issues. Our unique intersection of skills allows us to precisely measure and continuously optimize how our mapping architecture impacts customer performance.

Partner with the best

As a Site Reliability Engineer, you will collaborate with cross-functional teams to optimize the performance, availability, and reliability of Akamai’s core Mapping Service. In this high-impact role, you will define critical KPIs, advance our monitoring and alerting infrastructure, and architect automated operational responses. Operating at the intersection of systems engineering and data science, you will apply statistical analysis and cutting-edge machine learning to diagnose and solve the internet’s most complex content delivery challenges.

Note: This is a highly strategic engineering role requiring deep, independent analytical skills—not a standard QA, DevOps, or operational position.

As a Senior Site Reliability Engineer, you will be responsible for:

  • Drive Observability: Co-design, manage, and track product SLIs/SLOs to proactively monitor, investigate, and analyze system performance and availability.
  • Leverage Data Insights: Apply advanced analytical skills and statistical insights to identify mapping bottlenecks, resolve reliability challenges, and engineer long-term solutions.
  • Innovate Tooling: Build and deploy internal tools that automate proactive performance tracking and accelerate independent incident diagnosis.
  • Advocate for Reliability: Partner with product engineers to champion scalable, resilient, and highly supportable system architectures.
  • Influence Strategy: Provide data-driven insights to guide executive-level decision-making and identify high-impact technology investments.
  • Resolve Incidents: Collaborate with internal engineering teams to swiftly troubleshoot, root-cause, and resolve complex customer escalations.

Do what you love

To be successful in this role you will:

  • Master’s or PhD in Computer Science or a highly analytical equivalent field.
  • 5+ years of experience in Site Reliability Engineering (SRE) or a related engineering role.
  • Deep mastery of Unix/Linux internals, computer networking protocols, and distributed system design.
  • Professional fluency operating within a command-line UNIX/Linux computing environment.
  • Data & Observability

  • Strong background in statistical data analysis, SQL database querying, and data integrity troubleshooting.
  • Proven ability to transform complex datasets into actionable strategic roadmaps.
  • Practical knowledge of enterprise observability, logging, and alerting systems like Grafana. 
  • Software Engineering & Leadership

  • Coding proficiency in a major backend or scripting language (e.g., Python).
  • Self-motivated communicator capable of articulating complex systems to non-technical stakeholders while managing multiple timelines.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Similar Jobs

8 Days Ago
Remote
Shri Bhrigukshetra, BLR, Uttar Pradesh, IND
Senior level
Senior level
Fintech • Analytics
Design, build, and maintain reliable infrastructure platforms; drive automation with Python, Ansible, and Terraform; implement CI/CD and IaC; enhance observability, monitoring, and alerting; lead incident response, RCA, and service restoration; apply ITIL practices for incident, change, and problem management.
Top Skills: AlertingAnsibleAWSAzureCi/CdClickhouseCriblGrafanaInfrastructure As Code (Iac)MonitoringObservabilityOpentelemetryPythonTerraform
15 Days Ago
In-Office or Remote
India
Senior level
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Operate and improve the reliability, availability, and performance of large-scale GeForce NOW services. Participate in incident triage and on-call rotations, build automation and tooling, enhance observability (metrics/logs/traces), drive SLO/SRI practices, run postmortems, and design/operate Kubernetes-based services across cloud and datacenter environments.
Top Skills: AWSAzureBashContainerizationElk/OpensearchGCPGoGrafanaKubernetesMicroservicesOpentelemetryPrometheusPython
16 Days Ago
In-Office or Remote
India
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
Design, implement, and maintain reliable, scalable infrastructure for large distributed content delivery systems. Define and measure SLIs/SLOs, monitor availability and performance, troubleshoot incidents, and implement corrective actions. Develop automation to reduce manual work, participate in design reviews, and collaborate with product and engineering teams to improve system reliability and performance.
Top Skills: AdbmsBashCloud ComputingDatadogGrafanaJavaScriptOracle SqlPrometheusPythonUnix/Linux

What you need to know about the Kolkata Tech Scene

When considering the industries shaping India's tech scene, gaming might not immediately come to mind. However, in the last decade, increased internet usage and greater access to mobile devices have catapulted the industry to new heights, with Kolkata-based companies like Virtualinfocom, Red Apple Technologies and Digitoonz, at the forefront, driving the design and animation of new gaming titles for players.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account