Junior Site Reliability Engineer
We are seeking a motivated and detail-oriented OnCall Site Reliability Engineer (SRE) to join our dynamic team. The successful candidate will be responsible for maintaining the health and performance of our systems, ensuring high availability, and quickly responding to incidents. This role requires an understanding of our technical stack and the ability to work effectively under pressure.
Requirements
- 6+ months of experience in DevOps or system administrator;
- Understanding of cloud infrastructure and container orchestration;
- Familiarity with monitoring and logging tools;
- Strong problem-solving skills and attention to detail;
- Ability to work effectively in a team environment;
- Excellent communication skills;
- Native Ukrainian speaker.
Responsibilities
- Monitor and maintain the health of our systems, ensuring high availability and performance;
- Respond to incidents and troubleshoot issues in a timely manner;
- Collaborate with development and operations teams to implement improvements and optimize system performance;
- Create and maintain documentation for incident response and system maintenance procedures;
- Participate in on-call rotations to provide 24/7 support.
Technical Stack
- AWS;
- Kubernetes;
- Terraform;
- ElasticSearch;
- Kafka;
- Grafana;