Site Reliability Engineer (f/m/d)
experienced
full-time
You analyze incoming requests, narrow down issues in a structured way, and get to the technical root cause of problems. To do this, you work with logs, metrics, deployments, and configurations, using our monitoring setup to identify causes as quickly and reliably as possible. At the same time, you stay close to our customers, keep them updated on progress, and make sure solutions are documented in a clear and traceable way.
We operate and support our Kubernetes environments via kloudease, our managed Kubernetes platform.
In general, you should be willing to take on on-call duty. If you are on call and respond to an incident, both duties are, of course, compensated separately and can give your monthly salary a noticeable boost.
- Technical Support: You handle incoming tickets related to our Kubernetes environments, prioritize requests, and keep our customers transparently informed about the current status
- Troubleshooting & Debugging: You analyze technical issues using logs, metrics, deployments, and configurations and systematically narrow down the root cause
- Monitoring & Alerting: You work with Prometheus and other monitoring solutions, evaluate metrics, and use alerts specifically to support troubleshooting
- GitOps: You work with GitLab as well as GitOps-based setups using Argo CD or Flux CD and support the analysis of deployment and configuration issues
- Documentation & Communication: You document root causes and solutions in a clear and traceable way and translate technical findings into clear next steps – both internally and externally
- You have hands-on experience operating and supporting Kubernetes environments
- You enjoy working with customers and communicate clearly, professionally, and solution-orientedly, even when things get stressful
- You are confident in troubleshooting and debugging and can analyze logs, metrics, and system states in a structured way
- You have experience with GitOps, GitLab, and tools such as Argo CD or Flux CD
- You are familiar with monitoring and alerting, particularly with Prometheus
- Ideally, you have experience with Traefik or comparable ingress controllers
- You work in a structured way, prioritize tickets effectively, and stay on top of even more complex technical issues
- You communicate confidently in both German and English and can explain technical topics clearly to customers as well as internally
- Your life, your plan: Flexible and self-organized work in an agile corporate culture without hierarchical levels. Remote or in the office.
- Attractive additional benefits such as job bike, discounted Deutschland-Ticket (Public Transportation), corporate benefits, working abroad, sabbatical, company pension plan, comprehensive health package for your physical and mental well-being (EGYM Wellpass, Plus Card, OpenUp)
- Efficiency meets comfort: Customizable work environment with state-of-the-art hardware and software
- No standing still: €2,000 personal training budget and AOE Academy
- Together as a team, we do it all: summer party, team events, after-work get-togethers, internal FIFA league, board game night, movie night
- Coolest office in Wiesbaden: Right in the city center, free parking, Thai massage, our own chef, play area (foosball, table tennis, billiards), fruit, snacks, drinks






