At Mindera, We are seeking an experienced and highly skilled Site Reliability Engineer (SRE) / Azure Monitoring Engineer with 6 to 12 years of hands-on experience to join our dynamic team in Chennai (Hybrid). The ideal candidate will have a strong background in managing cloud infrastructure, automation, and monitoring, with a focus on Azure cloud services and modern DevOps practices. This role is crucial to ensuring the reliability, scalability, and performance of our systems using Azure monitoring tools and practices like GitOps.RequirementsKey Responsibilities:Azure Infrastructure Monitoring & Optimization:Develop and implement Azure monitoring solutions using tools like Azure Monitor, Application Insights, and Log Analytics to ensure the health and performance of cloud-based resources. Monitor and analyze system logs, metrics, and alerts from Azure services to detect and resolve issues proactivelyProficiency in logging, monitoring & alerting setups of Azure Experience in on call support and usage of on call support toolsKubernetes & Docker Management:Manage Kubernetes clusters and containerized applications using Azure Kubernetes Service (AKS)Implement and maintain containerization best practices with Docker, ensuring optimal performance of containerized workloadsIncident Management & Troubleshooting:Lead and manage incident response for performance, availability, and security issues within cloud infrastructureTroubleshoot and resolve issues related to Azure services, Kubernetes, containers, and VMs, ensuring rapid resolution and minimal downtimeAutomation & GitOps Implementation:Implement GitOps practices using GitHub Actions to automate deployment, monitoring, and incident management processesAutomate routine monitoring and infrastructure management tasks with PowerShell and other scripting languagesReliability Engineering:Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs) to ensure the reliability and uptime of servicesContinuously enhance the resilience of cloud infrastructure, applications, and services to meet high standards of performance and reliabilityCollaboration and Documentation:Collaborate closely with cross-functional teams, including DevOps, development, and infrastructure teams, to ensure integrated monitoring and automationDocument monitoring setups, procedures, troubleshooting guides, and incident response strategiesSecurity & Compliance:Ensure monitoring and automation practices are in compliance with internal security policies and best practicesPerform audits and implement security controls to safeguard cloud infrastructure and sensitive dataRequired Skills and Experience:Azure Cloud Services:Extensive experience with Azure services, including Azure Monitor, Application Insights, Log Analytics, and other monitoring solutionsKubernetes & Docker:Strong experience with Kubernetes and managing containerized applications in Azure Kubernetes Service (AKS)Proficient in Docker for containerization and managing containerized environmentsOperating Systems:Experience with Windows and Linux administration in cloud environments, including deployment, configuration, and troubleshootingScripting & Automation:Expertise in PowerShell and other scripting languages to automate monitoring and cloud infrastructure management tasks.GitOps & CI/CD:Hands-on experience with GitOps workflows and GitHub Actions for automating deployment pipelines and operational processesIncident Management & Troubleshooting:Proven experience in incident response, troubleshooting, and resolving cloud-related infrastructure issues, ensuring rapid recovery and minimal service disruptionReliability Engineering & Monitoring:Experience setting and managing SLOs, SLIs, and SLAs to measure and ensure system reliability and availability.Problem-Solving & Analytical Skills:Strong analytical and problem-solving skills to identify performance issues, their root causes, and to implement improvementsExperience In Scaled ApplicationsThe candidate should have worked previously in scaled programmes/services that cater to high volumes of concurrent users and also demands high availabilityCommunication SkillsDemonstrated experience in collaborating with clients across Europe and globally, comprehending various business requirements, and providing solutions that adhere to international standardsEngagements with multiple vendors - able to manage interactions with different partiesPreferred Qualifications:Azure Certifications:Azure certifications (e.g., Azure Administrator, Azure Solutions Architect, Azure DevOps Engineer) are highly preferredContainerization & Orchestration Tools:Familiarity with Helm for managing Kubernetes applications and other container orchestration tools.BenefitsWe OfferFlexible working hours (self-managed)Competitive salaryAnnual bonus, subject to company performanceAccess to Udemy online training and opportunities to learn and grow within the roleAbout M