Last week, at the @SupportVectorsL Weekly Paper reading Meetup (https://t.co/i5cPlkcjoD), I introduced Grokking in neural networks and its possible interpretations and ended with the GrokFast paper.
The presence of grokking in neural networks is one of those delightful and unsolved mysteries about deep learning, where the models develop unsuspected abilities. In this case, it goes against the grain of the classical statistical learning notion of bias-variance tradeoff: making more complex models on limited data, one should end up with a pathologically over-fitted model that has relied chiefly on internal memorization of the training data. However, with neural nets, if we continue to iterate over the learning loop for a few more orders of magnitude, then surprisingly, at some point, the models "grok" or learn to generalize correctly, and the validation loss falls rapidly. This phenomenon is called grokking, and researchers have speculated that it is related to another related phenomenon: double dipping.
SupportVectors AI training will offer a free, 3-session course in data wrangling. It starts August 1 (Thursday evening) at 7 PM PST.
Beginner's Introduction to Data Wrangling with Python https://t.co/8F78Mnw4No #Meetup via @Meetup
SupportVectors is pleased to announce a free 3-session course starting Thursday evening (August 1) at 7 PM PST.
Beginner's Introduction to Data Wrangling with Python https://t.co/6SeShBpS5T #Meetup via @Meetup
(Meetup paper reading session starts now)
Srikanth Chandar and Chandar Lakshminarayan will walk us through a reading of this 2022 paper on time-series forecasting with transformers:
A Time Series is Worth 64 Words: Long-term Forecasting with Transformers
----
Discord Server: https://t.co/EdUCPjufoN
We are a group of applied AI practitioners and enthusiasts who have formed a collective learning community. Every week, on Wednesday evening PM PST, we hold our research paper reading seminar, covering a topic in AI. One member walks through the paper with careful explanations, making it more accessible to a wider audience. Then we follow this reading with a more informal discussion, and socializing.
You are welcome to join this in person or over Zoom. SupportVectors is an AI training lab located in Fremont, CA, close to Tesla and easily accessible by road and BART. We follow the weekly sessions with snacks, soft drinks, and informal discussions.
If you would like to attend by Zoom, the Zoom registration link is below:
https://t.co/Qm4NPuX3ie
Last week, at the @SupportVectorsL Weekly Paper reading Meetup (https://t.co/i5cPlkcjoD), I introduced Grokking in neural networks and its possible interpretations and ended with the GrokFast paper.
The presence of grokking in neural networks is one of those delightful and unsolved mysteries about deep learning, where the models develop unsuspected abilities. In this case, it goes against the grain of the classical statistical learning notion of bias-variance tradeoff: making more complex models on limited data, one should end up with a pathologically over-fitted model that has relied chiefly on internal memorization of the training data. However, with neural nets, if we continue to iterate over the learning loop for a few more orders of magnitude, then surprisingly, at some point, the models "grok" or learn to generalize correctly, and the validation loss falls rapidly. This phenomenon is called grokking, and researchers have speculated that it is related to another related phenomenon: double dipping.
On a personal note, I recollect with fondness the first graduate course I took on the mathematical methods of classical mechanics. Our textbook was authored by V.I. Arnold -- and I was captivated by his insightful treatment of an old subject using manifolds and tangent bundles! The textbook is still available these days: https://t.co/Z5NlLf0iM3
Since KAN (Kolmogorov-Arnold Network) has been gaining a lot of research interest lately, one of the questions that keep coming up in discussions is: Why Splines? While there have emerged alternative implementations of KAN with other basis functions, such as Chebychev polynomials and wavelets, Splines in a way formed a natural first choice owing to their prior use in piecewise/local construction of complex functions.
As I explained the basics of splines, I thought it is worth mentioning a wonderful youtube video on the splines in depth, with beautiful visualizations and lucid explanations. https://t.co/DW4yibGR9c
The Eigen Layer allocation checker is now live! 🪂🔥
Quickly check your eligibility and potential rewards by visiting: https://t.co/PxYVoyp4uH
Simply connect your wallet to get started.
Don't miss out on this opportunity to claim your share of $EIGEN tokens. Mark your calendars for May 10th when claiming officially begins!
How many $EIGEN tokens did you qualify for? Let us know in the comments below!
👇🏻 #EigenLayer #Airdrop #CryptoCommunity $MOCA $LAVA $AVAIL $GRASS $DOGE $PIXEL $DIONE $TON $KLAY $BAG $METIS $KUJI $APT $CANTO $CRO $GMD $INJ $EGLD $ALGO $ALEO $STRK $CHKN $BVM $ONDO $APU $LINK #ChainSwap #chainlink $CSWAP $OPSEC $AMPL $FRAX #BasedAI $BASED #CSWAP $ENA $MUMU $DEAI $GFI $YAI #BASEDAI $GPU $ZKML $TD #CCIP $TD #CCIP #CCTP $PEAS $INFRA
The Eigen Layer allocation checker is now live! 🪂🔥
Quickly check your eligibility and potential rewards by visiting: https://t.co/PxYVoyp4uH
Simply connect your wallet to get started.
Don't miss out on this opportunity to claim your share of $EIGEN tokens. Mark your calendars for May 10th when claiming officially begins!
How many $EIGEN tokens did you qualify for? Let us know in the comments below!
👇🏻 #EigenLayer #Airdrop #CryptoCommunity $MOCA $LAVA $AVAIL $GRASS $DOGE $PIXEL $DIONE $TON $KLAY $BAG $METIS $KUJI $APT $CANTO $CRO $GMD $INJ $EGLD $ALGO $ALEO $STRK $CHKN $BVM $ONDO $APU $LINK #ChainSwap #chainlink $CSWAP $OPSEC $AMPL $FRAX #BasedAI $BASED #CSWAP $ENA $MUMU $DEAI $GFI $YAI #BASEDAI $GPU $ZKML $TD #CCIP $TD #CCIP #CCTP $PEAS $INFRA
Today, I am speaking at #TiEcon 2023 on the panel for Generative AI's impact.
At SupportVectors, we have been pondering over the societal impact of AI, and in particular, how it will reshape data engineering -- with a lively deba…https://t.co/pAXLywJSlK https://t.co/Kg4rJq80Jd
A dissenting perspective may be needed here. The article makes an argument that is true if and only if we can get away with transfer learning or if one of the existing ML algorithms works sufficiently well for the problem.
In other words, it presumes th…https://t.co/U5XPjx6xYX
This is a well-thought-out gem of an article that delves into the pragmatics of making the head and tail of these LLMs for real-life use cases. It captures the crux of the struggles when it says:
"LLM limitations are exacerbated by a lack of engineering…https://t.co/b09KyyZryi
Are we indeed seeing "Sparks of Artificial General Intelligence"?
In a remarkably long Microsoft research paper that showed up on https://t.co/cTZKMTFOFl this week, the authors present conversational evidence with GPT-4, which makes them wonder if GPT-4…https://t.co/rTGVayi6pO