Production Support & Incident Response Lead

  • Full-time

Company Description

Why EviSmart

  • 300 people. Two hubs: Vancouver HQ and Manila operations.
  • 145% year-over-year SaaS growth - the market is responding.
  • 28 countries. One platform. The dental industry's Autopilot.
  • An in-house AI model research and development team building proprietary intelligence.


How We Work • We ship before we're 100% certain. We write things down because we have two offices and memory is lossy. We debate loudly and move without resentment. We treat the customer's real problem as more important than an elegant internal process. If you've spent time waiting for permission to try something obvious — you'll notice the difference here immediately.

Job Description

Incident Commander | SaaS Production Operations
Permanent night shift • Individual Contributor • Full onsite – BGC, Taguig

This is not an escalation coordinator role. We need someone who can make real-time production recovery decisions when systems fail and customers are already affected.

When an incident happens, your job is not:
Follow procedure → identify resolver → escalate → wait → send updates.

Your job is:
Understand what is actually affected → determine the fastest safe recovery path → make the call → direct the teams involved → keep the customer operational → stay on it until production is restored.

What this actually looks like

EviSmart is a live SaaS platform with multiple applications, services, dependencies and operational workflows.

When something breaks, the answer will not always be written in a runbook.

You need to be able to determine: 

  • What is actually failing?
  • What is the customer impact?
  • Which systems or dependencies are involved?
  • Is the normal recovery path working?
  • How long should recovery reasonably take?
  • Can Production/Application Support restore it?
  • Do we actually need Engineering?
  • Is there a workaround or manual process that keeps the customer operational?
  • At what point do we stop waiting and change the recovery strategy?

Different failures have different recovery paths. During the incident, you own the clock and the decision.

If the normal recovery route isn't working, we expect you to recognize that quickly and change course rather than continue following a process that is no longer protecting the customer.

That can mean restoring the application, rolling something back, implementing a safe workaround, directing Support or Engineering, or switching operations to a manual process while the underlying issue is being resolved.

You stay accountable until production is operational again.

 

What we mean by Incident Commander

You command both the people and the recovery strategy.

Engineering may perform the permanent technical fix. Production Support may execute part of the recovery. Customer Support may need to communicate a workaround.

But during the incident, you are responsible for determining what needs to happen next and driving everyone toward recovery.

You need enough technical depth to investigate production issues yourself using things such as logs, monitoring, application behavior, databases, APIs, recent releases and system dependencies.

You also need enough judgment to know when to continue investigating, when to restore, when to escalate technically, and when the safest decision is to keep the customer moving another way.

 

Incident Management experience alone is not enough

You may have years of Incident Management experience and still find this role difficult.

Our production environment is a large, interconnected ecosystem. An issue in one part of the system can affect another application, workflow, team or customer process.

If most of your experience has been following established procedures, running bridges, coordinating resolver groups and escalating technical decisions to someone else, this is probably not the right role for you.

We need someone who can reason through an unfamiliar production problem even when:

the documentation is incomplete, the normal process isn't working, information is still coming in, teams disagree on the cause, and the person you would normally escalate to isn't available.

You won't know our ecosystem on Day 1. We can teach you that.

What is much harder to teach is the ability to analyze a complex system, make a defensible decision with incomplete information, take responsibility for that decision, and change course when the evidence tells you you're wrong.

If your background isn't a perfect match but you genuinely believe you can demonstrate that level of technical reasoning and judgment to us during the interview, we're willing to give you the opportunity to prove it.
 

We're likely to be interested if you come from

Production Support, Application Support, SaaS Operations, Site Reliability, Technical Operations, Incident Response, or another environment where you have personally handled business-critical production failures.

We're particularly interested in people who have worked with multiple interconnected applications or services and have personally made recovery decisions during live incidents.

 

The simplest test: 

At 2 AM, production is down, the runbook isn't solving it, Engineering isn't available, and customers need to keep operating.

Can we trust you to figure out what needs to happen next, make the call, and own the outcome?

If yes, we want to talk to you.

Qualifications

Ideally, you have:

  • Strong hands-on experience in Application Support, Production Support, SaaS Operations or a similar production environment

  • Experience supporting systems in a 24/7 or on-call environment

  • Real production incident troubleshooting experience

  • Experience with monitoring and observability tools such as Grafana, Datadog, CloudWatch, Splunk, Kibana, Azure Monitor or similar

  • Working knowledge of logs, APIs, SQL/databases, cloud environments and basic diagnostic tools

  • Experience with incident response, root-cause analysis and production releases

  • Strong judgment around escalation, risk and business continuity

  • The ability to communicate confidently with both technical teams and business stakeholders

 

Deep DevOps expertise is not required. What matters is that you can investigate intelligently, understand what you're seeing, take the safest action available at your level, and know when specialist intervention is genuinely necessary.


Important before you apply

This is a permanent night-shift role.

This is also an Individual Contributor role, not a traditional people-management position. You will act as the functional point person during night coverage and will help develop a designated backup, but you will not be joining to manage a large team.

The responsibility is significant because you are the person we need to trust when the daytime team isn't around.

If you're already strong in Application or Production Support and you're looking for an opportunity where you're given more ownership, more decision-making authority and the chance to become the person trusted to lead production incidents, we'd like to hear from you.

Apply and tell us about the toughest production problem you've personally solved.

Additional Information

🌟 Work setup: FULL ONSITE, Monday to Friday
📍 Location: One World Place, BGC, Taguig City
💼 Employment Type: Full-time + Permanent

 

Get to know us more:
EviSmart 
• https://www.instagram.com/evidentdigital/
• https://evismart.com/#work-flow-video 

By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply