Attentions concatenated
So, just started learning regarding Multi-head attention e penso que, com a explicação do Karpathy, consegui compreender o core do que seria o multi-head attention. Interesting because it decreases the embedding per se. Like, it was considering 32 for self-attention. With the consideration for concatenation, the number is 32, and each self-attention produces just a part and then they are concatenated.
Simple concept, but really considering the gain with this strategy.
Really great learning.