SigNoz Logo

SigNoz

Sr Site Reliability Engineer

Reposted 21 Days Ago
Remote
Hiring Remotely in India
Senior level
Remote
Hiring Remotely in India
Senior level
Own reliability, scalability, and operability of a petabyte-scale observability SaaS. Improve SLOs/SLIs, incident response, and on-call practices; scale and tune ingest pipelines and ClickHouse; manage Kubernetes clusters, autoscaling, multi-tenancy, and upgrades; build infra-as-code, CI/CD, capacity planning, and observability for the platform.
The summary above was generated by AI

About SigNoz

SigNoz is an open-source observability platform that helps modern engineering teams monitor, debug, and optimize their applications with deep visibility into metrics, traces, and logs — all in one place. We're built natively on OpenTelemetry and offer both self-hosted and cloud options, so teams can run observability the way they want, without vendor lock-in.

We are growing fast and building core developer infra products. And we are not fooling around:

  • 27,000+ GitHub stars

  • 800+ customers

  • 7,000+ members in our Slack community

Role: Sr Site Reliability Engineer (SRE)

We're looking for an SRE to own the reliability, scalability, and operability of the SigNoz cloud platform. You'll keep a petabyte-scale observability system fast and dependable — making sure the people who trust us to watch their systems can always trust ours. The platform team handles infra, scalability of SaaS, ingest pipelines, staging environments, automation, and the operational backbone of the product.

This is a deeply hands-on role for someone who understands what actually breaks in production at scale — and enjoys fixing it for good.

What we're looking for

  • Kubernetes at scale — not just "I've deployed to k8s," but real fluency with the nuances and gotchas: resource tuning, autoscaling behavior, networking, stateful workloads, upgrades, and the failure modes that only show up under load

  • Working knowledge of ClickHouse — operating it, tuning queries, and understanding its behavior at scale — is a strong plus

  • Knowledge of Golang is a plus (most of our stack and tooling is in Go)

  • Familiarity with OpenTelemetry and running large-scale data ingest pipelines is a plus

What you'll work on

You'll work with a high-caliber team across areas like:

  • Reliability of the SigNoz cloud platform: SLOs/SLIs, error budgets, incident response, and on-call practices that don't burn people out

  • Scaling the ingest path — making it robust to bursts while maintaining data freshness

  • SaaS auto-scalability and capacity planning across a petabyte-scale system

  • Operating and tuning ClickHouse and the data layer for performance and cost

  • Kubernetes infrastructure: cluster operations, upgrades, multi-tenancy, and the automation that keeps it boring

  • Observability of SigNoz itself — we dogfood our own product, so you'll help make it world-class

  • Infrastructure-as-code, CI/CD, and the tooling that lets a small team operate big systems

What will make you successful

  • 5–8 years in SRE, infrastructure, or platform/backend roles operating production systems at scale

  • Deep, practical Kubernetes experience — you know where the bodies are buried

  • Strong grasp of distributed systems failure modes, performance debugging, and capacity planning

  • Comfortable in code (Go preferred) — you automate and fix things, not just configure them

  • Loves open source — ideally with prior contributions to OSS projects (any size)

  • Comfortable in a high-ownership, fast-moving, remote-first environment

  • Strong communication — can write clear runbooks and tech docs and explain trade-offs

Nice-to-haves

  • Past experience on platform/infra/SRE teams of Series B+ startups

  • Hands-on experience operating ClickHouse, Kafka, or similar high-throughput data systems

  • Experience in observability (monitoring / logging / tracing) and with OpenTelemetry

Why you'll love working at SigNoz

  • Work on a globally used open-source project that engineers actually love

  • Huge scope and ownership — your work directly shapes how teams adopt SigNoz

  • Collaborate with a high-caliber team who just can't stop shipping

  • Remote-first, async-friendly culture

  • Opportunity to help define the future of open-source observability

Similar Jobs

2 Days Ago
In-Office or Remote
India
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
The Senior Site Reliability Engineer improves the reliability, scalability, availability, and performance of distributed content delivery systems. Responsibilities include defining SLOs and SLIs, monitoring platforms, debugging incidents, implementing corrective actions, automating operational processes, participating in design reviews, and guiding scalable infrastructure design. The role collaborates with Product and Engineering teams and applies software engineering, systems administration, cloud, DevOps, and SRE practices.
Top Skills: AdbmsBashCloud ComputingDatadogDevOpsGrafanaJavaScriptOracle SqlPrometheusPythonUnix/Linux
2 Days Ago
In-Office or Remote
India
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
Design, deploy, and maintain Akamai’s Compute platform infrastructure and internal tools. Improve service availability, reliability, scalability, observability, and performance through automation, monitoring, and configuration management. Support incident response and on-call operations, troubleshoot customer-impacting issues, define reliability requirements, and mentor other SRE engineers while collaborating across teams.
Top Skills: AnsibleBashDockerEnvoyGoGrafanaHaproxyJenkinsLinuxLokiNginxPrometheusPythonRedisSaltstackTerraform
11 Days Ago
In-Office or Remote
India
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
Oversee, scale, and optimize high-density AI hardware infrastructure across regional data centers. Build Python automation and infrastructure-as-code tooling, integrate incident workflows, develop telemetry pipelines and monitoring dashboards, and improve reliability across private cloud, bare-metal, and virtualized environments. Lead on-call incident response, runbooks, post-mortems, service rollouts, vendor coordination, and field technician activities while driving uptime, performance, and operational readiness.
Top Skills: Ai-Based Anomaly DetectionApi IntegrationsBare-Metal InfrastructureBgpGrafanaInfrastructure As CodeIpv4Ipv6LlmsLokiOpentelemetryPagerdutyPrivate CloudPrometheusPythonSlackTelemetry PipelinesVirtualization

What you need to know about the Delhi Tech Scene

Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account