Every entry below is a crawlable link to a unique public task page. This directory complements the interactive filters and keeps the complete catalog connected through standard pagination.
Tech Screening · Easy · 3 min
Users report slow connection setup, but application latency is normal. How would you separate TCP, TLS, and network delay?
View taskTech Screening · Medium · 3 min
A certificate rotation succeeds in one environment and fails in another. Walk through the trust chain you would inspect.
View taskTech Screening · Hard · 3 min
How would you set timeouts and retries across a service chain without creating retry amplification?
View taskTech Screening · Easy · 3 min
A downstream slowdown causes request volume to triple. What immediate controls would you apply, and what would you redesign?
View taskTech Screening · Medium · 3 min
One backend is overloaded while peers are idle. Which load-balancer and application signals would you inspect?
View taskTech Screening · Hard · 3 min
How would sticky sessions change scaling, failure recovery, and deployment behavior?
View taskTech Screening · Easy · 3 min
What makes a container image reproducible and trustworthy enough for production?
View taskTech Screening · Medium · 3 min
The same image tag points to different contents across clusters. How would you repair the release process?
View taskTech Screening · Hard · 3 min
A new pod enters CrashLoopBackOff. Which evidence would you inspect before editing the manifest?
View taskTech Screening · Easy · 3 min
Previous container logs show a missing secret, while readiness checks also fail. Which issue is causal, and how do you prove it?
View taskTech Screening · Medium · 3 min
Explain how startup, readiness, and liveness probes should differ for a slow-starting service.
View taskTech Screening · Hard · 3 min
An aggressive liveness probe turns a dependency slowdown into a restart storm. How would you redesign it?
View taskTech Screening · Easy · 3 min
How do CPU and memory requests and limits affect scheduling, throttling, and eviction?
View taskTech Screening · Medium · 3 min
Pods are Pending although cluster utilization looks low. What scheduler evidence would you inspect?
View taskTech Screening · Hard · 3 min
A new network policy blocks only DNS from one namespace. How would you verify the policy and restore service safely?
View taskTech Screening · Easy · 3 min
How would you introduce default-deny network policy without breaking unknown production dependencies?
View taskTech Screening · Medium · 3 min
A credential is committed to a repository. What must happen beyond deleting the file?
View taskTech Screening · Hard · 3 min
Design secret delivery and rotation for workloads running across several clusters.
View taskTech Screening · Easy · 3 min
When would you choose rolling, blue-green, or canary deployment, and what rollback signal matters most?
View taskTech Screening · Medium · 3 min
A canary is healthy at five percent but fails at fifty percent. Which capacity and traffic effects would you investigate?
View taskTech Screening · Hard · 3 min
A pipeline passes after rerun without a code change. How would you determine whether the test, cache, runner, or dependency is nondeterministic?
View taskTech Screening · Easy · 3 min
A release job is green but publishes an older artifact. Walk through the evidence you would compare.
View taskTech Screening · Medium · 3 min
How would you prove that a production image was built from the reviewed commit and expected dependencies?
View taskTech Screening · Hard · 3 min
What release metadata would let an incident responder trace a running binary back to source and build inputs?
View taskTech Screening · Easy · 3 min
How would you design a CI cache key so it speeds builds without reusing incompatible dependencies?
View taskTech Screening · Medium · 3 min
A cache makes builds fast but occasionally wrong. Which inputs, restore rules, and validation would you change?
View taskTech Screening · Hard · 3 min
A Terraform plan replaces a production database unexpectedly. What would you inspect before approving it?
View taskTech Screening · Easy · 3 min
Which changes in an infrastructure plan should trigger a hard approval gate rather than an ordinary review?
View taskTech Screening · Medium · 3 min
Terraform reports drift after an emergency console change. How would you reconcile state without hiding the cause?
View taskTech Screening · Hard · 3 min
What controls would you put around remote state, locking, access, and recovery?
View taskTech Screening · Easy · 3 min
A workload role has broad administrator access. How would you reduce privilege without causing an outage?
View taskTech Screening · Medium · 3 min
Design access for engineers who occasionally need production debugging privileges.
View taskTech Screening · Hard · 3 min
A private service is reachable from one subnet but not another. How would you trace routing, policy, and name resolution?
View taskTech Screening · Easy · 3 min
How would you separate security groups, network ACLs, routes, and application listeners during a connectivity incident?
View taskTech Screening · Medium · 3 min
Replica lag rises during a schema migration. How would you protect reads and decide whether to pause?
View taskTech Screening · Hard · 3 min
A database failover restores availability but changes connection behavior. What would you verify before closing the incident?
View taskTech Screening · Easy · 3 min
A queue backlog grows while consumers appear healthy. Which throughput, age, retry, and dependency signals would you inspect?
View taskTech Screening · Medium · 3 min
How would you drain a large backlog without overwhelming the downstream database?
View taskTech Screening · Hard · 3 min
A new metric label causes monitoring cost and query latency to spike. How would you contain and redesign it?
View taskTech Screening · Easy · 3 min
Which dimensions belong in metrics, and which are better kept in logs or traces?
View taskTech Screening · Medium · 3 min
When would you start an incident investigation with metrics, logs, or traces, and how do you connect them?
View taskTech Screening · Hard · 3 min
A latency regression appears only in traces. What observability gap allowed the customer impact to escape alerting?
View taskTech Screening · Easy · 3 min
How would you define a user-centered availability SLI for an API with retries and several endpoint classes?
View taskTech Screening · Medium · 3 min
A service meets its uptime target while checkout users still fail. What is wrong with the reliability measurement?
View taskTech Screening · Hard · 3 min
A team gets frequent CPU pages but misses customer-visible failures. How would you redesign the alert?
View taskTech Screening · Easy · 3 min
What makes an alert worthy of waking someone, and what context should it provide immediately?
View taskTech Screening · Medium · 3 min
How would you use an error budget to resolve conflict between release speed and reliability?
View taskTech Screening · Hard · 3 min
A team exhausts its error budget early in the month. Which actions should follow, and which should not be automatic?
View task