L'économie africaine 2025 : Les grandes tendances macroéconomiques en Afrique from Grandes lignes : le podcast de la recherche sur le développement de l'AFD 💬 Listen now on https://t.co/5RdV3iMxok
ou sur https://t.co/dWktjjeXFa
#IA#Afrique
🌟 The African Society in Digital Sciences presents the ASDS Webinar Series! 🌟
🛡️ Topic: "Privacy-preserving Synthetic Data Generation in Finance Using Agent-based Modeling for Fraud Detection."
📅 Date: October 25, 2024, at 4 PM (WAT)
🎯 L’objectif du DevFest ?
Réunir les développeurs, experts et passionnés de tech pour partager, apprendre et innover ensemble.
https://t.co/oDgbD6QzK7 🚀
#DevFest#DevFestYaounde#GDGYaounde
🎤 [Appel aux speakers]
Vous êtes expert(e) en technologie ? Participez au #DevFest Yaoundé 2024 en tant que speaker !
Partagez vos connaissances avec une audience passionnée.
Inscrivez-vous maintenant ! 🔗 https://t.co/3mppJbke5T
#DevFest#DevFestYaounde#GDGYaounde
Finding quality OCR datasets was a huge challenge 🤔
And then I see that one of the largest OCR datasets are now available to the public as open @huggingface dataset 🔥
With over 26 million pages , 18 billion text tokens, and 6TB of data. 🤯
These resources are just goldmine for document AI research.
Cosine-Similarity of Embeddings may not be always about similarity 🤔
📌 This Netflix paper cautions against blindly using cosine similarity and proposes alternatives such as training the model directly with cosine similarity, projecting embeddings back to the original space before applying cosine similarity, or applying normalization/popularity bias reduction before or during training.
✨ It concludes that cosine similarity can be arbitrary and meaningless depending on the regularization used during training.
It says cosine similarity can actually yield arbitrary and sometimes non-unique results depending on the regularization used when training the matrix factorization (MF) model
📌 Experiments on simulated data with known ground-truth item clusters illustrate the large variability in item-item cosine similarities for the first model under different choices of D, compared to the unique solution of the second model.
📌 Alternatives to Cosine-Similarity: To address the limitations, alternatives like layer normalization or avoiding the embedding space altogether are suggested. Applying normalization during or before learning can also enhance the semantic similarities, as seen in methods like word2vec's negative sampling.