Job Description
Summary
We are looking for an experienced Site Reliability Engineer (SRE) to join our fast-growing Engineering and R&D team. As a critical part of our infrastructure, you will ensure the reliability, scalability, and performance of large-scale distributed systems and applications. If you have strong expertise in Linux performance tuning, DNS architectures, TCP/L7 proxies, Kubernetes, automation, and CI/CD pipelines, this Site Reliability Engineer role offers a challenging opportunity to shape robust systems that deliver seamless user experiences at scale.
Job Details at a Glance
| Job Title | Site Reliability Engineer (SRE) |
|---|---|
| Location | India (Remote/Hybrid options available) |
| Employment Type | Full-time |
| Department | Engineering & IT / Engineering and R&D |
| Key Skills & Tools | Linux, Kubernetes, NGINX, HAProxy, DNSSEC, DHCP, BIND, Ansible, Puppet, Chef, Prometheus, Grafana, Splunk, Jenkins, GitHub CI/CD, Python, Golang, Shell scripting |
| Experience Level | Mid to Senior Site Reliability Engineer |
| Shift | Standard Business Hours (On-call rotation required) |
Responsibilities
As a Site Reliability Engineer (SRE), you will:
- Design, build, and maintain highly reliable, scalable, and efficient large-scale systems and applications.
- Optimize system performance through CPU thread management, memory tuning, and kernel optimizations.
- Architect and manage scalable DNS infrastructure using DNSSEC, DHCP, and BIND protocols.
- Configure and maintain TCP/L7 proxy services such as NGINX, HAProxy, and Squid for efficient traffic and content delivery.
- Automate DNS and proxy configurations using configuration management tools like Ansible, Puppet, or Chef.
- Monitor system health and performance with Prometheus, Grafana, Telegraf, and Splunk, proactively identifying bottlenecks and improvements.
- Deploy and manage containerized applications using Kubernetes for efficient orchestration.
- Troubleshoot complex multi-layer issues across CDN, load balancers, storage, OS, network, Kubernetes, virtualization, and application/database stacks.
- Build, maintain, and optimize CI/CD pipelines leveraging Jenkins and GitHub CI/CD to streamline software delivery.
- Develop automation scripts in Python, Shell, or Golang to reduce manual operational tasks and increase reliability.
- Collaborate closely with development teams to design scalable systems aligned with business goals.
- Participate in 24/7 on-call rotations to ensure rapid resolution of critical incidents and maintain high system uptime.
Skills & Qualifications
To succeed as a Site Reliability Engineer, you should have:
- Bachelorβs degree in Computer Science, Computer Engineering, or equivalent experience.
- Strong proficiency in Linux fundamentals including kernel tuning and performance optimizations.
- Deep knowledge of DNS protocols (DNSSEC, DHCP, BIND) and experience configuring/managing TCP/L7 proxies such as NGINX, HAProxy, and Squid.
- Hands-on experience with Kubernetes container orchestration and management.
- Proficient programming and scripting skills in Python, Shell scripting, and Golang.
- Expertise with configuration management and automation tools like Ansible, Puppet, or Chef.
- Solid experience with CI/CD tools such as Jenkins and GitHub Actions for pipeline setup and maintenance.
- Excellent problem-solving skills to diagnose and resolve complex multi-layer infrastructure issues.
- Familiarity with monitoring and alerting tools including Prometheus, Grafana, Telegraf, and Splunk.
- Strong communication skills and ability to collaborate across cross-functional teams.
- Experience in a 24/7 operational environment, including on-call support.
Why Join Us?
Join a cutting-edge team where your expertise as a Site Reliability Engineer will directly impact the stability and scalability of mission-critical systems. Work in a collaborative, innovative environment that values automation, efficiency, and continuous improvement. Help us deliver seamless, reliable services that empower users worldwide.
Apply Now
Ready to take your Site Reliability Engineer career to the next level? Join us to build resilient systems that matter.