
Om Jaju
DevOps and SRE Career Coach Resume Audits Interview Guides and Coaching
Habilidades

Revisa mis servicios


Experiencia laboral
Site Reliability Engineer
Morgan Stanley • Tiempo completo
Apr 2025 - May 2026 • 1 yr 1 mo
• Increased monitoring coverage from 30% → 99% by implementing Prometheus-based monitoring, alerting, and health checks across 7,560+ critical production systems, collaborating with stakeholders to establish monitoring best practices and ensure alignment with business-critical SLOs. • Eliminated 120+ recurring production alerts through root cause analysis (RCA), remediation, and reliability improvements, coordinating with client teams to validate fixes and ensure minimal business impact during remediation. • Reduced incident resolution time (TTR) by 60%+ by authoring 79+ SOPs and runbooks, standardizing troubleshooting workflows and partnering with stakeholder teams to validate runbook effectiveness in real-world incident scenarios. • Automated directory provisioning and access management using Python, eliminating 16–20 manual operational requests/month while coordinating with security and infrastructure stakeholders to enforce least-privilege access controls.