Your area of work:
We are looking for an experienced DevOps Engineer with strong expertise in Kubernetes, observability, and cloud-native operations. In this role, you will design, implement, and scale enterprise-grade monitoring and logging solutions across modern distributed systems running on Google Kubernetes Engine (GKE).
Beyond technical excellence, you will act as a trusted advisor and mentor, helping engineering teams adopt observability best practices, improve operational resilience, and leverage AI-driven solutions to enhance efficiency and reliability. You will play a key role in shaping our observability strategy and fostering a strong DevOps and Site Reliability Engineering culture across the organization.
Your responsibilities:
- Design, implement, and maintain scalable monitoring and logging solutions using Prometheus, Grafana, and Loki.
- Manage, optimize, and support workloads running on Google Kubernetes Engine (GKE).
- Build and continuously improve observability platforms that provide actionable insights into system performance and reliability.
- Ensure the availability, scalability, and resilience of monitoring and logging infrastructure.
- Implement and enhance CI/CD pipelines and automation processes for infrastructure and observability platforms.
- Apply Infrastructure as Code principles using tools such as Terraform and Helm.
- Integrate monitoring and observability practices throughout the software development lifecycle.
- Drive Site Reliability Engineering (SRE) practices, including SLIs, SLOs, alerting strategies, and incident response processes.
- Act as a technical mentor and subject matter expert for engineering teams.
- Lead workshops, knowledge-sharing sessions, and best practice initiatives related to observability and platform engineering.
- Enable teams to adopt self-service observability capabilities and standardized monitoring approaches.
- Explore and implement AI-powered tools and agents to automate operational tasks, improve incident management, and optimize monitoring processes.
- Contribute to continuous innovation within DevOps and platform engineering practices.
Your profile:
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Proven experience as a DevOps Engineer, Platform Engineer, or similar role in cloud-native environments.
- Strong hands-on experience with Google Kubernetes Engine (GKE) and Kubernetes ecosystem technologies.
- Deep expertise in Prometheus, Grafana, and Loki, including the design and operation of self-managed observability platforms.
- Solid understanding of Kubernetes architecture, networking, monitoring, logging, and alerting concepts.
- Experience with Infrastructure as Code tools such as Terraform and Helm.
- Knowledge of modern CI/CD platforms, including GitHub Actions, Jenkins, GitLab CI, or similar technologies.
- Familiarity with cloud platforms, preferably Google Cloud Platform (GCP).
- Experience with Site Reliability Engineering (SRE) principles and incident management processes.
- Strong communication, stakeholder management, and collaboration skills.
- Ability to mentor engineers and drive the adoption of DevOps and observability best practices.
- Experience building enterprise-scale observability solutions is highly desirable.
- Exposure to AI-powered operational tools, AIOps, or machine learning-driven automation is an advantage.
- Relevant Kubernetes and/or cloud certifications are considered a plus.