Three shapes cover nearly every case: batch scoring, online request-response, and streaming prediction. Batch runs on a schedule and writes predictions into a table or cache. Online serving computes a prediction inside a request, under a latency budget. Streaming scores events as they arrive on a queue and pushes results downstream.
Batch wins when the prediction does not depend on something the user just did. Churn scores, nightly recommendations, lead ranking, and credit pre-approvals all fit. You get cheap hardware, no tail-latency worry, and easy retries when a run fails. Reading a precomputed row is a key lookup, so the serving path stays trivial.
Batch loses the moment freshness matters. A score computed overnight is stale for a session that started at lunchtime. The usual cost of getting this wrong is a recommender that ignores everything the user did in the current session.
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓