Role Overview
We are looking for a DevOps / SRE Engineer with experience in managing production infrastructure, deployment, observability and system reliability. This position will play a role in maintaining platform stability and performance, managing Kubernetes on AWS, building CI/CD pipelines, and handling incidents and continuous improvement with the Engineering team.
Primary Responsibilities
Provide SRE support, including following on-call schedules, handling incidents, conducting root-cause analysis, and following up postmortems.
Manage deployment and release to production.
Maintained system observability and alerting using Datadog, Grafana/Prometheus, and Elasticsearch/Kibana.
Built and maintained CI/CD pipelines using Jenkins and GitHub Actions.
Manage and scale Kubernetes clusters on AWS to ensure reliability, scalability, upgradeability, and cost efficiency.
Qualification
Experienced in incident response/on-call, troubleshooting production issues, and creating postmortems.
Have strong hands-on experience in managing AWS in a production environment.
Mastering scripting/programming such as Bash and Python.
Have experience using Kubernetes in production, including deployment, scaling, upgrading, debugging, and networking.
Experience using Helm for Kubernetes packaging and releases, including authoring, templating, and versioning.
Experienced in managing cluster autoscaling and capacity, especially Karpenter.
Master CI/CD using Jenkins and GitHub Actions, including pipeline/workflow as code and reusable templates.
Experienced using ArgoCD / GitOps for continuous delivery.
Has production experience with Istio, including mTLS, traffic management, routing, and troubleshooting.
Mastering Infrastructure-as-Code using the Terraform ecosystem (Terraform, Terragrunt, and Atlantis) for multi-environment provisioning, orchestration, and PR-based plan/apply workflow.
Experience using observability tools such as Datadog, Grafana/Prometheus, and ELK (Elasticsearch/Kibana).
Good understanding of Linux and networking basics, including VPC, Load Balancer, DNS, TLS, and TCP/IP.
Experience using HashiCorp Vault for secrets management, including policy, authentication method, and secret rotation.
Experience managing data and messaging systems in production such as RDS PostgreSQL, MongoDB, RabbitMQ, and MQTT, including monitoring, backup/recovery, performance, and troubleshooting.
Strong experience with GitHub-based workflows, including branching, pull request review, and release and versioning practices.
Familiar with Sonar/SonarQube and Dependabot integration or similar tools into CI/CD.
Share Vacancies
Value-added
Have experience using Azure, especially App Service or Container App.
Have experience building CI/CD for Flutter iOS.
Have an advanced understanding of Helm Charts.
Experience managing multi-cluster Kubernetes or cross-region production environments.
Apply Now
Send your best CV and portfolio now. We are waiting for you to become part of the big Labamu family.