Senior Site Reliability Engineer
About the role
The role involves managing large-scale distributed system software deployments in cloud or on-premise environments with a strong foundation in cloud management. The candidate should have experience working with observability tools like Prometheus and Grafana, as well as incident handling and debugging skills. Experience with AWS Cloud (preferred) including hands-on experience with AWS-CLI, orchestration and containerization like Kubernetes, containers, CI/CD practices, networking, Linux OS, and shell/python scripting is required.
Candidates will join an innovative team pushing engineering boundaries. Depending on the domains associated with this job, you will be expected to design clean APIs, write automated test cases, and participate in peer code reviews. We value developers who focus on performance optimization, fast load times, and simple, maintainable architectures.