Breno

The forward layer to think

So, Karpathy agora mostrou a estrutura da arquitetura e, que, depois de multi-head attention, segue, no código, uma layer de forward, simple com linear layer e ReLU.

Great explicação de Karpathy. Like, why this layer? Ele explica que, depois dos tokens terem se conhecido, individually, they think about what they know regarding the other tokens.

Really great learning.