Senior Cloud Platform Engineer - CL

  • Full-time

Company Description

Technology is our how. And people are our why. For over two decades, we have been harnessing technology to drive meaningful change.
 
By combining world-class engineering, industry expertise and a people-centric mindset, we consult and partner with leading brands from various industries to create dynamic platforms and intelligent digital experiences that drive innovation and transform businesses.
 
From prototype to real-world impact - be part of a global shift by doing work that matters.

Job Description

About the Role

You will be the primary architect, operator, and developer-enablement partner for the client’s AWS engineering platform. This is a hands-on role responsible for designing, securing, scaling, and supporting the infrastructure, CI/CD services, developer tools, and observability capabilities that allow engineering teams to deliver software reliably. You will work across AWS, EKS/Kubernetes, Linux, networking, identity, automation, and shared engineering tools while mentoring teams on cloud-native practices.

 

Responsibilities

● Design and implement secure, multi-account structures and landing zones across AWS.

● Manage and evolve multi-cloud IAM, SSO, and Role-Based Access Controls (RBAC) to ensure least-privilege access.

● Enforce tagging standards, resource hierarchies, and cost-optimization strategies (rightsizing, idle resource elimination) to maintain fiscal accountability.

● Lead the deployment, scaling, and management of Kubernetes clusters (EKS, GKE, or self-managed). Manage CNI plugins, ingress controllers, and service meshes (Istio/Linkerd).

● Administer and optimize Linux (Ubuntu, Amazon Linux, RHEL) and Windows Server environments, ensuring hardened configurations and automated patching.

● Manage the intersection of cloud services and traditional OS-level dependencies, including Active Directory integration and file system performance tuning.

● Develop and maintain modular templates using Terraform, CloudFormation, or Pulumi.

● Build, orchestrate, and support CI/CD pipelines using Jenkins and Bitbucket, including webhooks, Jenkins agents, build/test stages, artifact publishing, approvals, deployments, rollback, notifications, and troubleshooting. Maintain GitOps workflows using GitHub Actions, GitLab CI, Flux, or ArgoCD where applicable.

● Design security controls including encryption at rest/transit (KMS), VPC Service Controls, and audit logging to meet SOC2, HIPAA, or FedRAMP standards.

● Leverage AI-native development tools (e.g., Cursor, GitHub Copilot) and LLM-powered agents to accelerate Infrastructure-as-Code (IaC) authoring, automate complex root-cause analysis, and proactively optimize cloud utilization through predictive anomaly detection.

● Install, configure, administer, upgrade, and integrate shared engineering tools such as Jenkins, SonarQube, Polaris, and Trino. Configure authentication, permissions, databases or supporting services, backups, health checks, and CI/CD integrations; manage secure, role-based, auditable infrastructure access using StrongDM.

● Enable developers to onboard applications, create and manage jobs, consume standard pipeline templates, and troubleshoot build, test, scan, deployment, and environment issues.

● Provision and operate AWS infrastructure, EC2 instances, EKS clusters, worker nodes, networking, load balancers, storage, databases, and supporting services using Infrastructure as Code.

● Design and maintain SSO, IAM, Kubernetes RBAC, service accounts, secrets, and read/write access models with least-privilege controls and auditable access reviews.

● Define operational ownership for the engineering platform, including tool availability, capacity, patching, upgrades, backups, disaster recovery, certificates, plugin lifecycle, and documented runbooks.

● Establish monitoring and observability for infrastructure, EKS, CI/CD, shared tools, databases, and applications. Progress from availability monitoring and metrics to centralized logs, dashboards, tracing, SLI/SLO-based alerting, incident response, and root-cause analysis.

● Monitor service health indicators such as uptime, CPU/memory/disk, network, latency, errors, throughput, pod and node health, Jenkins queue and job failure rates, scan failures, Trino coordinator/worker health, database connectivity, storage capacity, certificates, security events, and access failures.

● Partner with developers and security teams to define onboarding standards, quality gates, vulnerability policies, deployment controls, alert thresholds, and platform documentation.

Qualifications

Qualifications

AWS specific (Must have)

Client-specific must-have capabilities

Jenkins administration and Jenkins file/Groovy pipeline development

Bitbucket webhooks, pull-request integration, and source-control workflows

SonarQube administration and CI/CD integration

Polaris or equivalent application-security scanning platform administration

Trino deployment, configuration, access control, and operational support

Observability using CloudWatch, Prometheus, Grafana, ELK/OpenSearch, Loki, OpenTelemetry, or equivalent tools

Developer enablement, platform onboarding, runbooks, and production troubleshooting

SSO, IAM, Kubernetes RBAC, read/write permissions, secrets, and auditability

  • AWS Organizations & Landing Zones
  • IAM, SSO, RBAC
  • VPCs, Transit Gateway, networking
  • EKS administration
  • Terraform
  • CloudFormation
  • GitOps (ArgoCD/Flux)
  • Linux administration
  • Kubernetes security
  • Cost optimization
  • Multi-account governance
  • SonarQube
  • StrongDM

GCP specific (nice to have)

  • GKE
  • VPC Service Controls
  • IAM
  • Organization hierarchy

Azure  specific (nice to have)

VNET, subscriptions

Experience You’ll Need

● Bachelor’s or Master’s degree in Computer Science, Information Technology, or a related field.

● 8+ years of experience in Cloud Engineering, SRE, or DevOps, with deep proficiency in AWS and/or GCP.

● Expert-level experience administering Linux (shell scripting, kernel tuning) and Windows Server (Active Directory, Group Policy, PowerShell).

● Proven track record of running production-grade Kubernetes workloads at scale, including experience with Helm and container security.

● Strong proficiency in Python, Go, or Bash for infrastructure automation and tool development.

● Solid understanding of VPC/VNet design, peering, Transit Gateways, and zero-trust security models.

Client scope note: Requires hands-on ability to install or deploy, configure, integrate, secure, monitor, upgrade, and troubleshoot the listed tools—not only consume pre-provisioned services.

Preferred Qualifications

● Experience designing monitoring and observability from foundational infrastructure health through metrics, logs, dashboards, distributed tracing, SLI/SLOs, alerting, incident response, and continuous improvement using Prometheus, Grafana, ELK/OpenSearch, CloudWatch, OpenTelemetry, or equivalent tools.

● Background in migrating legacy Windows/Linux monolithic applications into containerized microservices.

● Relevant certifications: AWS Solutions Architect Professional, Google Professional Cloud Architect, or Certified Kubernetes Administrator (CKA).

Additional Information

At Endava, we’re committed to creating an open, inclusive, and respectful environment where everyone feels safe, valued, and empowered to be their best. We welcome applications from people of all backgrounds, experiences, and perspectives—because we know that inclusive teams help us deliver smarter, more innovative solutions for our customers. Hiring decisions are based on merit, skills, qualifications, and potential. If you need adjustments or support during the recruitment process, please let us know.

By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply

Privacy Notice