โ† back to jobs
> job detail
D
๐Ÿ‘ฝOther

Analyst, Infrastructure SRE Engineer, Data Platform, Group Technology

Dbs ยท Singapore - East
// classified as
Other (Adjacent or hard to classify.)
posted
1d ago
location
Singapore - East
languages
bash, python
tools
docker, grafana, kafka
> stack
bashpythondockergrafanakafkakubernetessparkterraform
> description

Role Overview:
As an Infrastructure Site Reliability Engineer (SRE) specializing in Kubernetes and GitOps, you will be instrumental in ensuring the reliability, availability, and performance of our critical Kubernetes-based data platform. You will apply software engineering principles to operations, focusing on system resilience, automation of operational tasks, and proactive incident prevention. This role requires a strong understanding of infrastructure as code and a commitment to continuous improvement of our production systems.

Key Responsibilities:

  • Design, implement, and manage highly reliable and scalable Kubernetes clusters and underlying infrastructure to support our data platform.

  • Implement and optimize GitOps workflows using ArgoCD for declarative continuous deployment and infrastructure configuration management, ensuring consistency and auditability.

  • Develop and maintain robust CI/CD pipelines using tools like Jenkins, focusing on infrastructure changes and automated validation for reliability and efficiency.

  • Define, monitor, and report on Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to drive continuous improvement in system reliability.

  • Implement comprehensive monitoring, logging, and alerting solutions (e.g., Prometheus, Grafana, OpenTelemetry) to proactively detect and diagnose infrastructure and application issues.

  • Automate infrastructure provisioning, configuration, and management using Infrastructure as Code (IaC) tools such as Terraform, Helm, Kustomize, and Ansible.

  • Lead incident response, root cause analysis (RCA), and post-incident reviews for critical infrastructure outages, implementing preventative measures to enhance system resilience.

  • Collaborate with development, operations, and data engineering teams to embed reliability practices throughout the software development lifecycle.

  • Manage Bitbucket repositories, ensuring version control, code review, and best practices for infrastructure-as-code.

  • Continuously identify and eliminate toil through automation and process improvements.

Job Requirements

  • Bachelor's degree in Computer Science, Engineering, or related field.

  • Strong experience with Kubernetes and containerization technologies (Docker, Podman).

  • Proficiency in Git version control system and experience with Git branching strategies.

  • Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or CircleCI.

  • Familiarity with infrastructure as code tools like Terraform, Helm, Kustomize, and Ansible.

  • Solid understanding of networking, security, and monitoring concepts in Kubernetes environments.

  • Strong scripting and automation skills (Bash, Python, or similar).

  • Excellent problem-solving and communication skills.

  • Ability to work effectively in a fast-paced, collaborative environment.

Nice to Have:

  • Certification in Kubernetes (CKA, CKAD) or related technologies.

  • Experience with observability platforms like Prometheus, Grafana, Splunk, or OpenTelemetry.

  • Knowledge of data processing frameworks like Apache Spark, Kafka, or Flink.

Location:

DBS Asia Hub

Job:

Analytics, Technology

Schedule:

Regular

Employee Status:

Full time