Every entry below is a crawlable link to a unique public task page. This directory complements the interactive filters and keeps the complete catalog connected through standard pagination.
Tech Screening · Easy · 3 min
How would you plan capacity for a seasonal service when growth, failover, and deployment headroom all matter?
View taskTech Screening · Medium · 3 min
A service has enough average capacity but fails during regional failover. Which assumptions were missing?
View taskTech Screening · Hard · 3 min
CPU-based autoscaling reacts too late for a queue-driven workload. Which signal would you use instead?
View taskTech Screening · Easy · 3 min
Autoscaling creates oscillation during traffic bursts. How would you tune the control loop?
View taskTech Screening · Medium · 3 min
During a major outage, how would you separate investigation, mitigation, coordination, and communication responsibilities?
View taskTech Screening · Hard · 3 min
Two teams propose risky fixes at the same time. How would you keep the incident moving without losing control?
View taskTech Screening · Easy · 3 min
When can rollback be more dangerous than continuing forward, and what evidence changes the choice?
View taskTech Screening · Medium · 3 min
An application rollback is available, but the release included a database migration. How would you verify compatibility first?
View taskTech Screening · Hard · 3 min
What makes a postmortem useful rather than a timeline followed by vague action items?
View taskTech Screening · Easy · 3 min
How would you write a causal explanation without reducing an incident to one person or one bad change?
View taskTech Screening · Medium · 3 min
How would recovery time and recovery point objectives change a multi-region disaster-recovery design?
View taskTech Screening · Hard · 3 min
A disaster-recovery plan exists but has never been exercised. What would you test before trusting it?
View taskTech Screening · Easy · 3 min
Why is a successful backup job not evidence that data can be recovered?
View taskTech Screening · Medium · 3 min
Design a restore test that proves data integrity, application compatibility, access, and recovery time.
View taskTech Screening · Hard · 3 min
A critical base-image vulnerability is announced. How would you determine exposure and roll out remediation safely?
View taskTech Screening · Easy · 3 min
A security patch requires restarting a stateful production fleet. How would you balance urgency and availability?
View taskTech Screening · Medium · 3 min
Which repetitive operational work should be automated first, and how would you prove the automation is safer?
View taskTech Screening · Hard · 3 min
An on-call task is frequent but highly variable. What would you standardize before attempting full automation?
View taskTech Screening · Easy · 3 min
One region is partially failing and sending bad responses. How would you decide whether to drain traffic?
View taskTech Screening · Medium · 3 min
Design failover verification for a service whose data, queues, and external dependencies span regions.
View taskTech Screening · Hard · 3 min
A production credential may be exposed. How would you rotate it without causing an outage or leaving old access valid?
View taskTech Screening · Easy · 3 min
Design routine secret rotation for clients that cannot all restart simultaneously. Which overlap and verification rules are required?
View taskTech Screening · Medium · 3 min
How would you prove that a production container came from reviewed source and an uncompromised build pipeline?
View taskTech Screening · Hard · 3 min
A dependency is replaced after CI passes but before deployment. Which signing and provenance checks should block release?
View taskTech Screening · Easy · 3 min
A Kubernetes service repeatedly scales up and down while latency worsens. How would you diagnose the autoscaling feedback loop?
View taskTech Screening · Medium · 3 min
Which metric would you scale a queue consumer on, and how would stabilization, startup time, and downstream capacity shape it?
View taskTech Screening · Hard · 3 min
Queue depth grows while consumers remain healthy. Which arrival, service-time, retry, and poison-message evidence would you inspect?
View taskTech Screening · Easy · 3 min
How would you contain a saturated asynchronous pipeline without silently dropping critical work or overwhelming its dependency?
View taskTech Screening · Easy · 3 min
A model is strong offline but weak online. How would you prove whether training-serving skew is the cause?
View taskTech Screening · Medium · 3 min
Your online feature values differ from the training snapshot. What would you inspect first, and how would you prevent recurrence?
View taskTech Screening · Hard · 3 min
Cross-validation looks excellent, but production quality collapses. How would you investigate target leakage without guessing?
View taskTech Screening · Easy · 3 min
A churn feature is populated after cancellation. Explain why that is leakage and how you would rebuild the dataset safely.
View taskTech Screening · Medium · 3 min
When is a random train-test split misleading, and how would you design a time-based validation instead?
View taskTech Screening · Hard · 3 min
Your users appear in both training and validation data across weeks. How would you prevent entity and temporal leakage?
View taskTech Screening · Easy · 3 min
Fraud prevalence is 0.2 percent. Which metrics and sampling choices would you use, and what could mislead you?
View taskTech Screening · Medium · 3 min
A classifier reaches 99.8 percent accuracy on an imbalanced dataset. How would you show whether it has any product value?
View taskTech Screening · Hard · 3 min
Two models have the same AUC but very different probability calibration. When does that difference matter?
View taskTech Screening · Easy · 3 min
Risk scores are consistently overconfident for new users. How would you measure and correct the calibration problem?
View taskTech Screening · Medium · 3 min
How would you choose a production threshold when false positives and false negatives have very different costs?
View taskTech Screening · Hard · 3 min
A stakeholder asks for a higher recall target. What evidence would you request before moving the decision threshold?
View taskTech Screening · Easy · 3 min
A key feature distribution shifts after a partner API change. How would you decide whether to alert, degrade, or retrain?
View taskTech Screening · Medium · 3 min
Prediction quality drops only in one region while aggregate metrics look stable. Walk through your drift investigation.
View taskTech Screening · Hard · 3 min
A real-time feature is six hours stale, but the endpoint is healthy. How would you detect and contain the impact?
View taskTech Screening · Easy · 3 min
What freshness contract would you define for online features, and what should serving do when the contract is violated?
View taskTech Screening · Medium · 3 min
How would you keep batch training features and low-latency serving features semantically consistent?
View taskTech Screening · Hard · 3 min
An online feature store and warehouse disagree for the same entity and timestamp. How would you localize the fault?
View taskTech Screening · Easy · 3 min
What should an online model do when a critical feature is missing for five percent of requests?
View taskTech Screening · Medium · 3 min
A default feature value keeps inference available but silently biases one cohort. How would you redesign the fallback?
View task