AceStack AI

Tech Screening

Inference batching

GPU utilization is low while request latency is high. How would you test whether batching or queueing is misconfigured?

What this task practices

Inference batching is a tech screening interview exercise that trains prompt interpretation, explicit assumptions, a concrete response, and a clear explanation of tradeoffs. The catalog marks it as medium difficulty. It focuses on Tradeoffs, Estimation, Technical Reasoning. The signed-in workspace provides the tools for the round and evaluates the attempt against task-specific criteria. Reference solutions, hidden checks, evaluator instructions, and candidate work remain private.

Sign in to startThe workspace and evaluation open after sign in.