Senior Infrastructure Engineer, Data Compute Platform
- Full-time
Company Description
About Grab and Our Workplace
Grab is Southeast Asia's leading superapp. From getting your favourite meals delivered to helping you manage your finances and getting around town hassle-free, we've got your back with everything. In Grab, purpose gives us joy and habits build excellence, while harnessing the power of Technology and AI to deliver the mission of driving Southeast Asia forward by economically empowering everyone, with heart, hunger, honour, and humility.
Job Description
Get to Know the Team
The Data Compute Platform team is an important contributor to Grab's data ecosystem, allowing growth through democratization of data at scale. We build and operate Grab's data infrastructure and efficient platform that supports internal data processes and company-wide data lake access. Our tech stack uses industry-leading distributed compute engines like Apache Spark, Ray, Trino, and Starrocks, orchestrated by Airflow and Michelangelo and backed by AWS S3. Our evolving Data Lake storage architecture uses modern open-source formats like Apache Iceberg and Delta in addition to traditional Apache Hive Parquet tables.
Underneath these engines sits a large, multi-tenant Kubernetes and AWS Infrastructure that we own end to end. This infrastructure includes the clusters, the operators, the autoscaling, the networking, the identity and access model, the observability, and the cost controls. These components work together to keep the platform fast, reliable, and affordable for thousands of pipelines and queries every day.
Get to Know the Role
You will be an important contributor to the infrastructure layer of the Data Compute Platform. This layer consists of the Kubernetes, AWS, and infrastructure-as-code foundations. Spark, Ray, Trino, Starrocks, Airflow, and Michelangelo run on these foundations. You will design how compute is provisioned, scaled, secured, observed and paid for, and you will drive the reliability and cost-efficiency of the platform as it grows. As a senior engineer, you will take ownership of well-scoped infrastructure projects from design through rollout and operation. You will also contribute to the team's engineering and SRE standards. Additionally, you will support other engineers through code review and knowledge sharing. You will also explore new developments in the cloud-native and data infrastructure space and integrate them into our ecosystem to the benefit of the data community at Grab.
You will report to our Data Engineering Manager II, and you will based onsite in our Petaling Jaya office.
The Critical Tasks You Will Perform
- You will design, build, and operate the multi-tenant Kubernetes (EKS) platform. This platform runs Grab's workloads, including Spark, Ray, Trino, Starrocks, Airflow, and Michelangelo. Additionally, your responsibilities will include cluster lifecycle, node provisioning, and autoscaling, and scheduling and resource isolation.
- You will build Kubernetes operators and custom resources (kubebuilder / controller-runtime, Go) that automate the provisioning and lifecycle of compute engines and their tenants.
- You will lead the infrastructure-as-code (Terraform) and CI/CD (GitLab) that provision and change our AWS estate. This estate includes EKS, S3, IAM, RDS, VPC, and networking. You will drive it towards safe, reviewable, automated change.
- You will design and implement the platform's identity, access and security model across AWS IAM, Kubernetes RBAC and service identities, working with the storage access and security teams.
- You will build the observability, alerting, capacity planning, and incident tooling for the platform. You will contribute to the SRE practice, which includes SLOs, runbooks, on-call, and post-incident reviews. You will reduce toil and MTTR.
- You will own compute cost efficiency: instance and storage strategy, spot and right-sizing, bin-packing, idle reclamation, and cost attribution back to tenants.
- You will drive architectural improvements and migrations (for example engine version upgrades, cluster consolidation, new execution backends), managing the design, phased rollout and rollback plan with guidance from senior team members.
Qualifications
What Essential Skills You Will Need
- Software Engineering, Computer Science, or related undergraduate degree.
- You have 3 or more years of experience, with at least 2 years building and operating production infrastructure or platform services at scale.
- You have programming proficiency in Go and/or Python, with the habit of treating infrastructure as software: tested, reviewed, versioned and automated.
- You have deep, hands-on Kubernetes experience in production: cluster operations, scheduling, autoscaling, networking, storage, RBAC and multi-tenancy. We prefer experience building custom controllers or operators.
- You have experience with AWS (EKS, EC2, S3, IAM, VPC) and infrastructure as code with Terraform.
- You have proficiency in CI/CD tooling (GitLab CI, Jenkins or similar) and GitOps-style delivery.
- You have solid SRE fundamentals: observability (Datadog, Prometheus, Grafana or equivalent), SLOs, incident management and capacity planning, with experience running reliable services.
Skills that are Good to have
- Working knowledge of at least one of Spark, Ray, Airflow, Trino or Starrocks, and a appetite to learn how distributed data engines behave on Kubernetes.
- Experience running Apache Spark on Kubernetes at scale (Spark Operator, dynamic allocation, shuffle services) and tuning its interaction with the resource manager.
- Experience with Kubernetes autoscaling and scheduling tooling such as Karpenter, Cluster Autoscaler, Yunikorn or Volcano.
- Experience deploying and operating Ray on Kubernetes (KubeRay), including cluster autoscaling and isolation for ML and batch workloads.
- Experience operating distributed query engines like Trino or Starrocks, including cluster sizing, fault tolerance and workload isolation.
- Experience with container runtimes and internals (containerd, cgroups, OverlayFS, image build and distribution).
- Experience with service mesh, ingress and network policy (Istio, Envoy, Cilium or similar) in multi-tenant clusters.
- Experience with FinOps for compute platforms: cost attribution, and spot / reserved capacity strategy.
- Contributions to open-source cloud-native or data infrastructure projects.
Additional Information
Life at Grab
We care about your well-being at Grab, here are some of the global benefits we offer:
- We have your back with Term Life Insurance and comprehensive Medical Insurance.
- With GrabFlex, create a benefits package that suits your needs and aspirations.
- Celebrate moments that matter in life with loved ones through Parental and Birthday leave, and give back to your communities through Love-all-Serve-all (LASA) volunteering leave
- We have a confidential Grabber Assistance Programme to guide and uplift you and your loved ones through life's challenges.
- Balancing personal commitments and life's demands are made easier with our FlexWork arrangements such as differentiated hours
What We Stand For At Grab
We are committed to building an inclusive and equitable workplace that provides equal opportunity for Grabbers to grow and perform at their best. We consider all candidates fairly and equally regardless of nationality, ethnicity, race, religion, age, gender, family commitments, physical and mental impairments or disabilities, and other attributes that make them unique.
By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply