The DevOps loop is a series of outages, described out loud
Sergei Skrylkov, Founder, Trippi · facts checked against the product on 2026-09-30
Most DevOps and SRE loops are built around one skill: staying methodical while something is broken. The questions are scenarios — the site is slow, the pod keeps restarting, the disk is full — and the interviewer plays the system, answering what each of your commands would show. There is no single right answer. There is a right order, and the grade is mostly whether you keep it while talking.
The loop, round by round
Recruiter screen — 20–30 minutes
Cloud provider, orchestration, infrastructure-as-code tool, on-call experience, salary.
Technical screen — 45 minutes, an engineer
Quick questions across Linux, networking, containers and one cloud. Breadth, not depth. General mechanics on the technical interview page.
Troubleshooting scenario — 60 minutes, a senior engineer
“The site is down, go.” You say what you would run; they say what it shows. Sometimes on a real broken machine in a shared terminal.
Infrastructure design — 45–60 minutes
A CI/CD pipeline, a multi-region setup, a migration to containers. Close to the system design round, with more on rollout and failure.
Scripting — 45 minutes
A small Bash or Python task: parse a log, rotate files, call an API with retries.
Behavioral — 45 minutes, a manager
An outage you caused, how you handle on-call, a time you automated away a manual process.
Ten questions, and what each one tests
| Question | What it tests | What a strong answer does |
|---|---|---|
| “A user says the site is slow. What is your first command?” | Troubleshooting method | Narrows the scope first — everyone or one user, all pages or one — before touching servers. |
| “A pod is in CrashLoopBackOff. What do you check?” | Kubernetes in practice | Exit code and events, logs of the previous run, resource limits, probes. |
| “What happens when you run kubectl apply?” | The control-plane model | Desired state sent to the API server; controllers bring actual state towards it. |
| “Design a pipeline for twenty microservices.” | CI/CD at scale | Build once, promote the same artifact, shared templates, fast feedback per service. |
| “Blue-green or canary — when do you use each?” | Rollout risk | Canary for gradual exposure with metrics; blue-green for fast, full switch and rollback. |
| “How do you manage secrets?” | Security basics | A secret manager, short-lived credentials, rotation, and never in the repository. |
| “Where does your Terraform state live, and what goes wrong?” | Infrastructure-as-code in teams | Remote state with locking, drift from manual changes, how you detect it. |
| “SLI, SLO, SLA — what is the difference?” | Reliability language | Measurement, internal target, external promise — with an error budget example. |
| “The disk is 100% full on a production host at 3 a.m.” | Calm sequencing | Free space safely first, then find the cause — including deleted files still held open. |
| “Tell me about an outage you caused.” | Ownership, blameless culture | What you did, how it was found, the fix, and the guardrail added after. |
Worked example: the pod that keeps restarting
In a scenario round the interviewer wants a sequence, said as commands and reasons, and they will answer each step. Here is the kind of structure a card offers for the opening question — points to speak from, not a runbook to read.
Example of a card’s shape — written for this page, not a screenshot
The line it grew from
“The payments pod is in CrashLoopBackOff since the last deploy. Walk me through what you do.”
Points to speak from
- 1Read why it died: describe the pod for the exit code and events, then the logs of the previous container, not the current one.
- 2Exit code 137 usually means it was killed for memory — check the limits; a failing liveness probe can look the same.
- 3Since it started with a deploy, say how you would roll back first to stop the impact, and then compare image, config and secrets.
After this, the interviewer will say what the logs show. That line appears in the window too, and the next card starts from it.
If English is your second language: tool names and on-call words
| What you hear | What it is |
|---|---|
| “Kube-control”, “kube-cuttle”, “kube-C-T-L” | All the same tool: kubectl. Nobody agrees on the pronunciation. |
| “Kates” | Kubernetes, from the abbreviation k8s. |
| “Engine-X” | nginx. |
| “Toil” | Repetitive manual work that should be automated. |
| “Runbook”, “playbook” | Written steps for handling a known alert. |
| “Drift” | Real infrastructure no longer matching the code that describes it. |
| “Roll back” vs “roll forward” | Return to the previous version, or fix by deploying a new one. |
| “Error budget” | How much unreliability the SLO allows before releases slow down. |
Scenario rounds are fast and full of short replies — “it shows 137”, “memory is fine”, “nothing in the logs”. Each one changes your next step, and missing one sends you down the wrong branch for five minutes.
What Trippi Cue does in a DevOps loop, and what it cannot
Trippi Cue is a Chrome extension. One click on its icon in the call’s tab in Google Meet, Zoom on the web or Teams on the web, and it listens to that tab’s audio — not your microphone, and no bot joins. The interviewer’s short replies arrive as lines, translated if you want, so “exit code one-three-seven” is on screen as text. When a question is addressed to you, a card follows with two or three points built from the last minutes of the scenario, so each step follows from what the system just “showed”.
Limits, plainly: it does not see the shared terminal, the broken machine, the dashboards or the logs on screen. If the interviewer only points at a graph, Cue knows nothing about it; if they read the value aloud, it does. It does not hear your own voice and it runs in browser tabs only.
FAQ
DevOps engineer, specifically.
What is asked in a DevOps interview?+
Linux, networking, containers and Kubernetes, one cloud, CI/CD, infrastructure-as-code, monitoring, and troubleshooting scenarios where you talk through what you would check.
Is DevOps interviewing the same as SRE?+
Close. SRE loops add more on reliability targets, capacity and coding; DevOps loops add more on pipelines and tooling.
How do I answer a troubleshooting scenario?+
Narrow the scope first, say each command and why, and listen to what the interviewer says it shows. Keeping the order matters more than guessing the cause fast.
Can Trippi Cue see the terminal they share?+
No. It hears the call’s audio only. Values said out loud reach it; text on the shared screen does not.
Does it work in the Teams desktop app?+
No. Join from Teams on the web in Chrome or Edge; Teams asks once for permission.
Read next
Short replies, each one a clue. Keep them on screen.
Trippi Cue is in the Chrome Web Store. One click on the icon when the call begins, and nothing before that.
