
Site Reliability Engineer Job Openings in Bangalore 2026!!!
SolarWinds announced job vacancy for the post of Site Reliability Engineer.The place of posting will be at Bangalore.Candidates who have completed Graduate / Engineering / Post Graduate with Fresher / Experience are eligible to apply. More details about qualifications, job description and roles & responsibilities are as follows
Company Overview
| Name of the Company | SolarWinds |
| Required Qualifications | Graduate |
| Skills | Strong hands-on experience with AWS and Azure cloud infrastructure |
| Category | Technology |
| Work Type | Onsite |
SolarWinds is looking for a Senior Site Reliability Engineer with 2+ years of experience to help build, operate, and improve the systems that support their products and internal platforms. This role is ideal for someone who is hands-on, dependable, and comfortable working across Kubernetes, AWS, Azure, Linux, and Database environments in a production setting.
Job Details
Θ Positions: Site Reliability Engineer
Θ Job Location: Bangalore
Θ Salary: As per company standards
Θ Job Type: Full Time
Θ Requisition ID: 202846
Roles and Responsibilities:
- The right candidate has a foundation in site reliability engineering, enjoys solving problems in live environments, and understands the importance of reliability, automation, and operational discipline.
- Operate, maintain, upgrade and improve production Kubernetes clusters and workloads across AWS and Azure environments.
- Manage Kubernetes platform components and technologies such as Helm, Kustomize, operators, istio, autoscaling, and cluster/node lifecycle management.
- Support and maintain production database platforms such as ClickHouse, Aurora and other distributed data systems, including performance troubleshooting and operational health.
- Build and maintain infrastructure using Terraform.
- Develop automation and tooling to reduce operational toil and improve reliability using Python, Go, Bash, or similar technologies.
- Participate in an on-call rotation, respond to production incidents, and lead or contribute to incident resolution and root-cause analysis.
- Improve observability, monitoring, logging, alerting, and incident response across infrastructure and services.
- Partner closely with software engineering and platform teams to design and deploy reliable, scalable services.
- Contribute to capacity planning, performance optimization, patching, upgrades, and infrastructure lifecycle management.
- Develop and maintain operational documentation, runbooks, and troubleshooting guides.
- Participate in reliability initiatives, including automation, resilience engineering, disaster recovery, and continuous improvement.
Required Skills & Qualifications:
- 2+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Platform Engineering, or a related field.
- Strong hands-on production Kubernetes experience, including operating, upgrading and troubleshooting Kubernetes clusters and workloads.
- Strong understanding of Kubernetes fundamentals, including: Pods, Deployments, StatefulSets, DaemonSets, Jobs, CronJobs Services, Ingress, DNS, and Kubernetes networking ConfigMaps, Secrets, and persistent storage
- Resource requests/limits and scheduling
- Autoscaling across workloads, clusters, and nodes
- Pod Disruption Budgets and high availability patterns
- Kubernetes upgrades and cluster lifecycle management
- Experience troubleshooting Kubernetes at both the application workload and cluster infrastructure levels.
- Strong hands-on experience with AWS and Azure cloud infrastructure.
- Strong Linux systems administration and troubleshooting skills.
- Experience operating customer facing, highly available production systems.
- Experience participating in a production on-call rotation and responding to high-severity incidents.
- Strong experience with Terraform and Infrastructure as Code.
- Experience with scripting and automation using Python, Go, Bash, or similar languages.
- Strong understanding of infrastructure concepts including compute, networking, storage, DNS, load balancing, and security.
- Strong troubleshooting skills with the ability to methodically diagnose complex distributed-system failures.
- Strong communication skills and the ability to collaborate effectively with engineering and cross-functional teams.
What they are looking for
- Ownership mindset and strong operational discipline
- Ability to stay calm and effective during incidents
- Willingness to learn, improve systems, and drive reliability-focused changes
- Practical approach to solving infrastructure and operational problems
- Team player who values clear communication and documentation
- Continuous learner, stays current with Kubernetes, cloud infrastructure, and modern SRE practices.
How to Apply
Apply Link – Click Here
For Regular Updates Join our WhatsApp – Click Here
For Regular Updates Join our Telegram – Click Here
Disclaimer:
The information provided on this page is intended solely for informational purposes for Students, Freshers & Experience candidates. All the recruitment details are sourced directly from the official website and pages of the respective company. Latest MNC Jobs do not guarantee job placement, and the recruitment process will follow the company’s official rules and Human Resource guidelines. Latest MNC Jobs do not charge any fees for sharing job information. Latest MNC Jobs strongly advise Students, Freshers & Experience candidates not to make any payments for any job opportunities.