In calculus, going from a single variable to millions of variables is hard.
Understanding the three main types of functions helps make sense of multivariable calculus.
Surprisingly, they share a deep connection. Let's see why!
Inverse Scaling: When Bigger Isn’t Better
There are tasks for which increased model scale alone may not lead to progress, and that more careful thought needs to go into the data and objectives for training language models.
repo: https://t.co/ePGmC6mtl3
abs: https://t.co/QIi4UnuoYC
On October 31, 1903, Frank Nelson Cole took to the stage at a now famous meeting of the American Mathematical Society with the aim of unveiling the factors of the Mersenne number 2⁶⁷ - 1, denoted as M67. Édouard Lucas had earlier, in 1876, established that M67 was not prime and must indeed have factors, but had fallen short of identifying what those factors were.
In a dramatic display, Cole began his so-called "lecture" in silence. He approached the chalkboard and meticulously calculated the value of M67, the result being a staggering 147,573,952,589,676,412,927. Shifting to the opposite side of the board, Cole then penned 193,707,721 × 761,838,257,287 and embarked on the arduous task of manually performing the multiplication. Upon equating the result to M67, Cole returned to his seat, having maintained silence throughout the hour-long spectacle. His audience rewarded this feat with a standing ovation.
Cole later confessed that the identification of these factors was the fruit of "three years of Sundays."
@dcmiringkar@val_iisc@iiscbangalore@CVPR Context is 'attending globally-renowned conferences', not 'immigration'.
Tell university & industry recruiters not to use international conferences as a criterion for candidate-filtering.
One can do a lot in & for the country when one's family can be decently fed & housed.
One lesson for writing NumPy/JAX code that took me a surprisingly long time to learn is to preserve array dimensionality whenever possible (i.e., always use keepdims=True).
@molynj_twtr@timurkuran@shashj Pre & post OneChildPolicy effects rippling in every generation (20-30 years).
USSR has inverse phenomenon. Good % of young men died in WW2. Resulting lack of children & future generations shows up every 20-30 yrs in their population pyramids.
@TimDuy 32 is the most common age in the US right now. What does a large population age 25-35 mean for the economy?
-People form more households by getting rid of roommates
-Buy their first homes
-Move to the suburbs
These are core trends driving the housing market
@jennifernvictor@Noahpinion One child policy came in 1980. Folks who survived Great Leap Forward had more kids in 1979 (Group A kids) than 1980 (Group B kids). After 20-30 yrs, A's kids outnumbered B's. This spike will happen every generation (20-30 yrs).
Ex-USSR has inverse pattern due to men dying in WW2.
CancerGPT - proposes a few-shot learning approach based on LLMs to predict the synergy of drug pairs in rare tissues that lack structured data and features.
CancerGPT (~124M parameters) is comparable to a larger fine-tuned GPT-3 model (~175B) on drug pair synergy prediction.
https://t.co/Wgqs2j2HG6
@therohit1986@the_yushi@Noahpinion Population control? Consider status quo, age the populn in the pyramid by 20 yrs, and then estimate proportion of people in different age groups. There would be fewer adults, more senior citizens and much more middle-aged people than present. Ergo, populn control not needed.
A prematurely aging population is among Mao’s legacies. In the Industrial Era, it would have spelled disaster. Will it be as harmful in the nascent AI Era? By reducing the number of job seekers, this aging might have the unintended effect of stabilizing China politically.
There's a tradition of film directors and studios congratulating each other for beating their box office records. A THREAD
In 1977, when STAR WARS beat Jaws to become the highest-grossing movie ever, Steven Spielberg took out the below ad for George Lucas in
@Variety
1/11
Not sure how I missed it, but this is a great overview of efficient NLP methods.
While models are scaled for performance, there is also a need to increase efficiency in modern NLP models. This survey paper covers methods and findings in efficient NLP.
https://t.co/lZ9dyX6OUb