[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

  • Full-time

Company Description

We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!
 

About the Client

Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world. Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies.

You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production.

Contract Duration: Initial contract through the end of 2026, extending the engagement to a total 12-month term based on performance.

Job Description

About the Role

This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack.

This is not general infrastructure, and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations, observability, Node.js runtime troubleshooting, JVM / Java service troubleshooting, Splunk, and incident ownership.

What You’ll Do

  • Support the deployment, operation, and reliability of production services running on Kubernetes.
  • Monitor service health and investigate production incidents across distributed applications.
  • Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
  • Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.
  • Support CI/CD, GitOps-based deployments, observability, and production monitoring.
  • Work within a client-directed backlog and established priorities.

Qualifications

Required Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.
  • 3+ years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting
  • Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene
  • Splunk experience for log aggregation, search, and production troubleshooting
  • Prometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing dashboards
  • CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux
  • Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking
  • Node.js production troubleshooting, including heap snapshots, CPU profiles, event-loop blocking, memory growth, worker/process isolation, and V8 isolates or similar runtime models
  • JVM / Java production troubleshooting, including GC log analysis, thread dump analysis, JVM tuning, and Java service latency investigation
  • In-memory cache experience with Redis / Valkey, including key design, TTL / eviction tuning, and cache invalidation
  • Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication

Additional Information

Nice to Have

  • Web Components / Lit experience, to perform first-level debugging of UI-related issues
  • Server-side rendering or isomorphic runtime experience
  • Canary rollout / multi-version production operations
  • Distributed tracing and request-context correlation
  • KEDA or event-driven autoscaling
  • Experience with enterprise platform integration layers

What We Offer

  • Competitive salary and laptop
  • Professional development and training opportunities
  • Work with cutting-edge cloud and container technologies
  • Flexible work arrangements and collaborative team environment
  • Impact on organization-wide digital transformation initiatives

By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply

Privacy Notice