Platform Engineer

  • Full-time

Company Description

Blend is a premier AI services provider, committed to creating meaningful impact for its clients through the power of data science, AI, technology, and people. We help organisations solve complex business challenges by combining deep domain understanding with modern data and AI capabilities. Our teams work across strategy, analytics, engineering, and product delivery to create scalable, high-value solutions that improve decision-making, efficiency, and growth

Job Description

You’d be joining the software engineering team behind a production-grade, event-driven platform at the intersection of cloud infrastructure and AI. The system processes complex, multi-step workflows in near real-time with AI capabilities woven throughout — running on Azure, managed as a Python Nx monorepo.
This isn’t a role that values depth in one discipline above all else — it’s a role that values breadth. We’re looking for a mid-level engineer who has worked meaningfully across software engineering, data engineering, and DevOps, and who can use that perspective to understand how a complex distributed system behaves end-to-end. Your primary focus will be building and owning the observability layer for the platform: Grafana dashboards that give the engineering team clear visibility into system health, performance, and behaviour. You’ll work collaboratively with the team to discover what needs to be measured, build the tooling to surface it, and grow into the team’s subject matter expert on observability.

Responsibilities
• Design and build Grafana dashboards that monitor the platform across multiple dimensions: service health, throughput, latency, queue depths, error rates, and alerting for impending or ongoing issues
• Instrument and surface metrics from across the system — including AI-specific signals such as LLM token consumption and inference timing
• Work closely with the engineering team to identify the data points that matter and translate them into useful, actionable views
• Connect observability tooling to Azure Monitor and OpenTelemetry data sources
• Set up and maintain alerting to give the support team early warning of degradation or 
failure
• Act as the team’s SME on observability — helping engineers understand what the dashboards reveal, coaching them to build their own dashboard components, and growing collective capability across the team
• Contribute to the broader platform engineering effort across CI/CD, deployment 
tooling, and cloud infrastructure as needed

Qualifications

  • Meaningful experience across at least two of: software engineering, data 
  • engineering, and DevOps/platform engineering — breadth here is genuinely the 
  • point, not a compromise
  • • Ability to read and understand a complex distributed system and reason about what 
  • to measure and why
  • • Hands-on experience with Grafana, including building dashboards from real data 
  • sources
  • • Familiarity with Azure monitoring tooling — Azure Monitor, Application Insights, or 
  • equivalent
  • • Understanding of observability concepts: metrics, traces, logs, alerting, SLIs/SLOs
  • • Experience working with event-driven or microservices architectures — you need to 
  • understand how the system works to know what to watch
  • • Python skills sufficient to navigate and contribute to the codebase

By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply

Privacy NoticeImprint