SURENDRA KUMAR
Site Reliability Engineer and DevOps professional with 10+ years of industry experience spanning cloud and on-premises infrastructure, application production support, automation, and system reliability. Proven expertise in leading incident management, root cause analysis, and code-level debugging with development teams on Java/Gosu applications running on Guidewire PolicyCenter. Skilled in defining SLAs/SLOs, driving postmortems and preventative measures, tracking vulnerability and Checkmarx mitigation to secure applications, and using observability tools such as Splunk and Dynatrace to maintain high availability in production environments. Proficient in Python and PowerShell automation, RHEL 8 infrastructure management, and hybrid cloud environments (AWS, Azure). Experienced in stakeholder reporting, mentoring junior engineers, and embedding SRE best practices across financial services, insurance, telecom, and technology sectors.
- Working in the financial services and insurance sector, ensuring high availability and reliability of critical applications.
- Supporting and troubleshooting applications built on Java and Gosu within the Guidewire PolicyCenter platform.
- Analyzing error logs and collaborating with developers to identify and fix code-level issues, reducing recurring incidents.
- Using Dynatrace and Splunk in production for monitoring, alerting, log analysis, and incident triage.
- Leading incident management and tracking, driving timely resolution and conducting postmortems to identify preventative measures.
- Defining and maintaining SLAs and SLOs to drive accountability and measurable reliability improvements.
- Tracking application vulnerabilities and preparing monthly Checkmarx mitigation reports to ensure secure, smooth application operation.
- Conducting weekly and monthly review meetings with stakeholders to report on progress, reliability metrics, and open action items.
- Mentoring junior team members on troubleshooting techniques, monitoring tools, and SRE best practices.
- Automating operational tasks and supporting scripting needs using Python and PowerShell to improve efficiency and reduce manual effort.
- Managing and supporting infrastructure on RHEL 8 servers in a production environment.
- Conducting root cause analysis and performance tuning to enhance application and infrastructure stability.
- Collaborating with development and operations teams to embed SRE best practices into workflows and improve service reliability.
- Supporting hybrid environments, integrating cloud and on-prem infrastructure for seamless operations.
- Designed and maintained CI/CD pipelines to streamline application deployment and reduce release cycles.
- Automated infrastructure provisioning using Terraform and configuration management tools.
- Deployed and managed Dockerized applications, improving scalability and portability.
- Implemented monitoring and alerting systems to enhance reliability and reduce downtime.
- Collaborated with cross-functional teams to integrate DevOps practices into development workflows.
- Troubleshot complex Linux/Windows issues, ensuring smooth operations across hybrid environments.
- Designed and implemented new network infrastructure, ensuring scalability and reliability.
- Executed integration of network components into existing systems, enabling seamless service delivery.
- Coordinated end-to-end deployment activities, from design to "ready to use" rollout.
- Conducted testing, validation, and documentation of installations for compliance and operational readiness.
- Collaborated with cross-functional teams to optimize deployment timelines and minimize service disruption.
- Performed maintenance and upgrades of telecom infrastructure, ensuring network stability and service continuity.
- Coordinated with regional teams to execute infrastructure upgrades with minimal downtime.
- Conducted preventive maintenance and troubleshooting to reduce recurring faults.
- Supported rollout of new infrastructure components, aligning with operational standards.
- Documented maintenance activities and upgrade reports for compliance and audit purposes.
- Led installation and commissioning of new infrastructure for telecom and network systems.
- Coordinated with cross-functional teams to ensure timely project delivery and compliance with technical standards.
- Conducted site inspections, troubleshooting, and quality assurance during deployment phases.
- Documented installation procedures and prepared handover reports for client validation.
Don Bosco Institute of Technology (DBIT), Bengaluru | 2010–2014
- Azure AI Fundamentals (Sep 2025)
- Azure Fundamentals (May 2025)
- AWS Certified Cloud Practitioner (Dec 2024)