๐๐๐๐ ๐๐ ๐๐๐ ๐ ๐๐๐-๐ ๐๐๐๐๐๐ ๐๐๐๐๐๐๐ (๐ ๐ ๐) ๐๐ ๐ ๐๐๐๐๐๐ ๐๐๐๐๐?
DEFINITION
-> The Feed-Forward Network (FFN) is a neural network component inside each Transformer block
-> It applies additional non-linear transformations to the representations produced by the attention layer
-> Unlike self-attention, the FFN processes each token's representation independently
-> It helps the Transformer learn complex patterns and transform contextual information into richer features
WHERE THE FFN FITS IN A TRANSFORMER
-> Input tokens are converted into embeddings
-> Self-attention allows tokens to exchange contextual information
-> The resulting representations are passed into the FFN
-> The FFN transforms each token representation
-> The transformed representation continues to the next Transformer layer
BASIC FFN STRUCTURE
-> The standard FFN contains two learned linear transformations with a non-linear activation between them
-> Wโ expands the hidden representation
-> bโ is the first bias term
-> ฯ is a non-linear activation function
-> Wโ projects the expanded representation back to the model dimension
-> bโ is the second bias term
STEP 1: INPUT FROM ATTENTION
-> The attention mechanism produces a context-aware representation for every token
-> Each token now contains information influenced by other relevant tokens
STEP 2: EXPANSION
-> The first linear layer usually increases the dimensionality of each token representation
-> For example, a model dimension of 4,096 may be expanded to a much larger FFN dimension
-> This gives the network more capacity to learn complex transformations
STEP 3: NON-LINEAR ACTIVATION
-> An activation function introduces non-linearity
-> Without this non-linearity, multiple linear layers would effectively behave like a single linear transformation
COMMON ACTIVATIONS
-> ReLU
-> GELU
-> SwiGLU
-> GeGLU
Modern LLM architectures frequently use gated variants such as SwiGLU rather than a simple two-layer ReLU/GELU FFN.
STEP 4: PROJECTION
-> The second linear transformation reduces the expanded representation back to the model's hidden dimension
-> This produces the output that can be passed to the next part of the Transformer block
POSITION-WISE PROCESSING
-> The same FFN parameters are applied independently to every token position
-> Tokens do not directly communicate with each other inside the FFN
-> Contextual interaction happens primarily through the attention mechanism
FFN VS SELF-ATTENTION
SELF-ATTENTION
-> Allows tokens to interact with other tokens
-> Learns relationships between different positions
-> Produces context-aware representations
FEED-FORWARD NETWORK
-> Processes each token independently
-> Applies learned non-linear transformations
-> Builds richer internal representations
WHY FFNs ARE IMPORTANT
-> Add substantial learnable capacity to Transformer blocks
-> Capture complex non-linear patterns
-> Transform features extracted through attention
-> Often contain a large portion of the model's parameters
-> Help LLMs develop sophisticated internal representations
FFN IN A TRANSFORMER BLOCK
-> Input representation
-> Self-attention
-> Residual connection + normalization
-> Feed-forward network
-> Residual connection + normalization
-> Output representation
MODERN LLM FFNs
-> Modern LLMs often use gated feed-forward architectures
-> A gating mechanism controls how information flows through the hidden representation
-> SwiGLU is a widely used example
-> Gated FFNs can provide better modeling capacity than traditional activation-only FFNs
SIMPLE EXAMPLE
-> Suppose a token representation has 4,096 dimensions
-> The FFN expands it into a much larger intermediate representation
-> A non-linear or gated activation transforms the representation
-> The second projection returns it to 4,096 dimensions
-> The resulting representation is passed deeper into the Transformer
SIMPLE FFN FLOW
-> Context-Aware Token Representation
-> Linear Expansion
-> Non-Linear Activation
-> Linear Projection
-> Transformed Token Representation
IN SIMPLE TERMS
-> Self-attention decides which information is important from other tokens
-> The FFN then processes and transforms that information
-> Together, attention and FFNs form the computational core of a Transformer block
Grab the MODERN LLMS ebook:
https://t.co/ljEMt0UNUI
๐จJeu Concours #BlackFriday๐จ #TPMPJeu
RT+Follow @rueducommerce et @TPMP pour tenter de gagner une vitrine de 20 produits๐ #RDC
Informatique, TV, Son, Photo, Smartphone, Electromรฉnager... La totale pour un quotidien plus confortableโก๏ธhttps://t.co/PEEKN5SfP7
TAS le 29/11/21๐
L'orthographe du mot aoรปt, quelle merveille ! Des douze mois de l'annรฉe il est le seul qui porte un chapeau. Pour nous prรฉserver du soleil, bien รฉvidemment.