Gateway/proxy: the single entry point in front of all model calls. Handles auth, per-tenant rate limiting and quotas, request/response logging, caching, key management, and provider routing. It is the control plane that keeps model access from sprawling across your codebase.
Orchestration: the application logic that turns a user request into one or more model calls. Prompt assembly, retrieval, tool calling, multi-step workflows or agent loops, retries, and fallback logic live here.
Model layer: the LLMs themselves plus any embedding and reranking models, possibly spanning multiple providers, regions, and tiers.
Retrieval: vector stores, keyword indexes, and databases that inject relevant context (RAG) so the model answers from your data, not just its weights.
Memory/state: storage for conversation history, user profiles, and long-term agent memory, since the API is stateless.
Guardrails: input and output checks for safety, prompt injection, personal data, and format validation, placed around the model calls.
Cross-cutting concerns wrap all of it: observability (tracing, token accounting, cost), evals, and a feedback/data flywheel.
Rewriting in plainer words…
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.