Hyperparameters are the knobs you set before training starts, as opposed to weights, which the training learns. A handful matter far more than the rest.
Learning rate: the size of each weight update, and by far the most sensitive.
Batch size: examples per update, which also drives memory use and throughput.
Epochs: how many passes over the data, usually settled by early stopping.
Depth and width: layers and units per layer, which set the model's capacity.
Regularization strength: dropout probability and weight decay, trading fit against generalization.
Optimizer: plain descent, momentum or Adam, each with its own sensible defaults.
Tune them roughly in that order. Sweeping everything at once wastes compute, because learning rate alone can swamp the effect of every other choice.
Rewriting in plainer words…
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.