O que aconteceu no FlyPodcast não é uma coisa normal. Estamos a falar de 63,177 milhões de kzs em 24h, isso não é normal. Promovam o Fly, invistam numa causa maior, façam alguma coisa, mas impulsionem esse camarada. Esse governo me irrita 🙄
@MADRlDISTAS essa cultura de player power não existe em lugar algum, só no real madrid. eu nunca gostei das taticas do xabi alonso mas o que o vini provocou é uma red flag do tamanho de um lençol (mesmo se o team estivesse 100% funcional e ele fosse bola de ouro)
Free Palestine
Free Sudan
Free Congo
Free Haiti
Free Cuba
Free Venezuela
Free Puerto Rico
Free Hawaii
Free all colonized people
Free the world
Free them all!
Embeddings are everywhere in modern AI, but the word "embedding" is doing a lot of different jobs.
A token embedding is a learned row in a vocabulary matrix. A contextual embedding is a hidden-state vector whose value depends on the surrounding sequence. A sentence or document embedding usually requires another step => some pooling or model-specific readout that turns variable-length representations into one fixed-size vector. These are related ideas, but they are not interchangeable.
The geometry is learned too. There is nothing inherently semantic about a vector just because it has 768 or 1,024 dimensions. The training objective has to shape the space so that useful pairs score well and unrelated examples separate. That is why contrastive training, hard negatives, pooling strategy, normalization, and the similarity function matter so much in modern retrieval systems.
This also explains several production bugs that look harmless at first. Cosine similarity and dot product are not generally the same, though they coincide for unit-normalized vectors. Changing the embedding model while keeping the same dimensionality does not preserve the coordinate system. Ignoring query/document prefixes can change results for models trained asymmetrically. Changing preprocessing or pooling can invalidate an existing index even when every vector still has the expected shape.
I put together a technical handbook that works through embeddings from first principles => token lookup tables, contextual representations, pooling, cosine/dot/L2 geometry, contrastive learning, dense retrieval, hard negatives, Matryoshka representations, multilingual and multimodal embeddings, and the implementation details that matter when these vectors are used for search and RAG.
The mental model I kept coming back to while writing it is simple => an embedding is useful geometry learned for a purpose. To understand one, you need to know what was mapped, how the vector was produced, what objective shaped the space, and how those vectors are compared.
Sharing the handbook here:
@yr_spec@UncleTimmy9@newsambarii If thats the point you disagree you could just have replied to that tweet but do you understand that you comment to a completely different topic? The person started talking about hip hop in general and you started talking about my opinion on it (??). Literally just read it again
@yr_spec@UncleTimmy9@newsambarii its not about picking the angle, the whole thread has been about the state of music/hip hop, i stayed consistent to it, some other dude even said that my argument doesn’t apply to Drake and i didn’t disagree because its not my point. IM NOT TALKING ABOUT DRAKE