Challenges & Incidents
Two stories and the production process. This is the middle of the interview.
Situation → what I found → what I did → result. About 90 seconds. End with a number.
These stories are drafts — adapt them to something you actually touched, because they'll ask follow-up details.
Tell me about a technical challenge you faced.
Your default answer.
- Situation — the weekly schedule screen took about 12 seconds for the biggest site, around 900 agents. Smaller sites were fine, so nobody had noticed. The client escalated.
- What I found — the code loaded the agents, then queried each agent's shifts inside a loop. About 900 round trips. And the shift table had no index on the column being filtered.
- What I did — replaced the loop with one query that pulled agents and shifts together and returned only the columns the screen showed. Added a covering index. Added server-side paging.
- Result — 12 seconds to under 1. Escalation closed. I then found the same pattern on two other screens and fixed those too.
Agents were occasionally assigned the same shift twice. Service Bus delivers at least once, and our handler assumed it would only see a message once — so a retry created a duplicate. Fixed it by recording the message id in the same transaction as the work, so a repeat does nothing.
I measured rather than guessing. Application Insights showed which call was slow, then
turning on SQL logging showed it running 900 times, and the execution plan showed the missing
index. I don't change anything until I can explain why it's happening, otherwise you can't tell
whether your fix worked.
Reproduce → measure → narrow down. If nothing was deployed, look outside: data volume, a dependency, a clock change.
Number → business outcome → what you fixed beyond your ticket.
About 12 seconds down to under 1, and the escalation was closed. The part I was more pleased
with is that once I knew the pattern, I found it on two other screens and fixed those before
anyone reported them.
Tell me about a production incident you handled.
Impact first, then cause.
- Situation — at 7am the Australian team reported real-time adherence hadn't moved since about 1am. Managers were making staffing decisions off six-hour-old numbers that didn't look obviously wrong.
- What I found — a schema change released the day before added a required field. Messages from the older version couldn't be read, so the consumer failed, retried, dead-lettered, and everything behind it was stuck.
- What I did — first made the consumer tolerate the missing field so the queue drained and live data returned, about 40 minutes. Then replayed the dead-lettered messages to backfill the gap. Then versioned the message contract and added an alert on dead-letter depth.
- Result — six hours of stale data for one region, no data lost. The alert has caught two smaller issues since, both before anyone noticed.
First how bad it is and for how many people, because that decides how fast we move. Then what
changed — the last release is always the first suspect. Then I follow one failing request
through the logs until the story stops making sense.
Stabilise → communicate → diagnose → fix → prevent. Everyone starts at diagnose. Don't.
Work out the scope, then check whether we deployed recently. If a release lines up, I roll
back first and investigate afterwards — customers don't care why it broke while it's still
broken. And I keep people updated while that's happening, because silence is what turns a
technical problem into a relationship problem.
Never start by reading code.
Keep asking why until you reach something you can change — and that's almost never a person's name.
There are usually two causes: one let it happen, one let it go unnoticed. Say the word blameless — interviewers listen for it.
Every incident should leave behind a test that would have caught it and an alert that would
have told us sooner. For that one it was a test against an older message version, an alert on
dead-letter depth, and a rule about versioning contracts. All three went on the board as
tickets, because "we should really add an alert" never survives the next sprint.