- Location
- Hyderabad, IN
- Work mode
- On-site
- Employment
- Full-time
- Experience
- Mid-level
About the role
Key Responsibilities Operational Excellence & SRE Drive Site Reliability Engineering (SRE) practices, including SLIs, SLOs, SLAs, error budgets, and automation of operational tasks. Manage incident response, root cause analysis, and post-incident reviews to strengthen platform resilience. Build and optimize observability and monitoring frameworks (CloudWatch, Grafana, Loki, Tempo, Prometheus). Implement self-healing systems and automated recovery where possible. Oversee OS patching to ensure no…
Requirements
- sre
- incident response
- root cause analysis
- observability
- monitoring
- automation
Skills
- sre
- site reliability engineering
- incident response
- root cause analysis
- observability
- monitoring
- automation
- cloudwatch
- loki
- tempo
About the company
Unison Consulting
Posted via Adzuna
How to apply
Apply Now takes you to Rozgoo, where auto-apply can submit your application for this role. Updated 8 days ago.
Apply Now