A · Your Work → Module 02
Module 02 · Questions 6–13

Challenges & Incidents

Two stories and the production process. This is the middle of the interview.

Every story has the same four beats

Situation → what I found → what I did → result. About 90 seconds. End with a number.

These stories are drafts — adapt them to something you actually touched, because they'll ask follow-up details.

Q6

Tell me about a technical challenge you faced.

Story 1 — the 12-second screen

Your default answer.

  • Situation — the weekly schedule screen took about 12 seconds for the biggest site, around 900 agents. Smaller sites were fine, so nobody had noticed. The client escalated.
  • What I found — the code loaded the agents, then queried each agent's shifts inside a loop. About 900 round trips. And the shift table had no index on the column being filtered.
  • What I did — replaced the loop with one query that pulled agents and shifts together and returned only the columns the screen showed. Added a covering index. Added server-side paging.
  • Result — 12 seconds to under 1. Escalation closed. I then found the same pattern on two other screens and fixed those too.
Backup story, if they want reliability instead

Agents were occasionally assigned the same shift twice. Service Bus delivers at least once, and our handler assumed it would only see a message once — so a retry created a duplicate. Fixed it by recording the message id in the same transaction as the work, so a repeat does nothing.

Q7How did you investigate it?
Say this

I measured rather than guessing. Application Insights showed which call was slow, then turning on SQL logging showed it running 900 times, and the execution plan showed the missing index. I don't change anything until I can explain why it's happening, otherwise you can't tell whether your fix worked.

Reproduce → measure → narrow down. If nothing was deployed, look outside: data volume, a dependency, a clock change.

Q8What was the result?

Number → business outcome → what you fixed beyond your ticket.

Say this

About 12 seconds down to under 1, and the escalation was closed. The part I was more pleased with is that once I knew the pattern, I found it on two other screens and fixed those before anyone reported them.

Q9

Tell me about a production incident you handled.

Story 2 — the overnight queue backup

Impact first, then cause.

  • Situation — at 7am the Australian team reported real-time adherence hadn't moved since about 1am. Managers were making staffing decisions off six-hour-old numbers that didn't look obviously wrong.
  • What I found — a schema change released the day before added a required field. Messages from the older version couldn't be read, so the consumer failed, retried, dead-lettered, and everything behind it was stuck.
  • What I did — first made the consumer tolerate the missing field so the queue drained and live data returned, about 40 minutes. Then replayed the dead-lettered messages to backfill the gap. Then versioned the message contract and added an alert on dead-letter depth.
  • Result — six hours of stale data for one region, no data lost. The alert has caught two smaller issues since, both before anyone noticed.
Q10How do you troubleshoot a production issue?
1 What's broken, for whom, since when 2 What changed — the last deploy is the first suspect 3 Follow ONE failing request through the logs 4 Narrow to one layer: app, database, cache, dependency 5 Confirm before changing anything
Say this

First how bad it is and for how many people, because that decides how fast we move. Then what changed — the last release is always the first suspect. Then I follow one failing request through the logs until the story stops making sense.

Q11What do you do first when production is failing?
The order

Stabilise → communicate → diagnose → fix → prevent. Everyone starts at diagnose. Don't.

Say this

Work out the scope, then check whether we deployed recently. If a release lines up, I roll back first and investigate afterwards — customers don't care why it broke while it's still broken. And I keep people updated while that's happening, because silence is what turns a technical problem into a relationship problem.

Never start by reading code.

Q12How do you do root-cause analysis?

Keep asking why until you reach something you can change — and that's almost never a person's name.

Adherence was stale └ the consumer couldn't read a message └ a required field was added to the contract └ nothing tested a version mismatch ← cause 1 └ and nobody noticed for 6 hours └ nothing alerted on the dead-letter queue ← cause 2

There are usually two causes: one let it happen, one let it go unnoticed. Say the word blameless — interviewers listen for it.

Q13How do you stop it happening again?
Say this

Every incident should leave behind a test that would have caught it and an alert that would have told us sooner. For that one it was a test against an older message version, an alert on dead-letter depth, and a rule about versioning contracts. All three went on the board as tickets, because "we should really add an alert" never survives the next sprint.