In a nutshell
A transformer processes a whole sequence at once instead of word by word. In each block, multi-head attention mixes information across tokens, then a small feed-forward network refines each token on its own; residual shortcuts and normalization keep it stable. Stack many such blocks and you get models like GPT and BERT.