We first learned how important Twitter was to the Japanese from the movie Shin Godzilla.
The political thriller version of Godzilla has this great scene where the older boomer President is saying the monster will never come out on land, but everyone is tweeting video of Godzilla walking on land and his younger subordinates have to react quickly to the threat of shin Godzilla
it may sound foolish, but we really loved Shin Godzilla and that scene cemented how important Twitter was
all our life we have loved Godzilla films, but none have captured the real terror of Godzilla or given him new powers and abilities quite the way Shin Godzilla did, we never saw Godzilla shoot radiation from his tail or back before Shin Godzilla and we've seen all the Godzilla movies
Shin Godzilla was a masterpiece of absolute cinema and I hope my Japanese friends love Godzilla as much as we do ๐
Yes, #RocketLeague is getting rid of trading on 12/05/23.
Let's end it right. Giveaways on my Twitch channel!
Over 1,000 items available
โ Must like and repost
โ Must follow channel & be present to win
โฐ 11/11 - 11/18 - 11/24
3:00/6:00pmEST
https://t.co/2mG5loQdBT
Yes, #RocketLeague is getting rid of trading on 12/05/23.
Let's end it right. Giveaways on my Twitch channel!
Over 1,000 items available
โ Must like and repost
โ Must follow channel & be present to win
โฐ 11/11 - 11/18 - 11/24
3:00/6:00pmEST
https://t.co/2mG5loQdBT
GPT-4 is getting worse over time, not better.
Many people have reported noticing a significant degradation in the quality of the model responses, but so far, it was all anecdotal.
But now we know.
At least one study shows how the June version of GPT-4 is objectively worse than the version released in March on a few tasks.
The team evaluated the models using a dataset of 500 problems where the models had to figure out whether a given integer was prime. In March, GPT-4 answered correctly 488 of these questions. In June, it only got 12 correct answers.
From 97.6% success rate down to 2.4%!
But it gets worse!
The team used Chain-of-Thought to help the model reason:
"Is 17077 a prime number? Think step by step."
Chain-of-Thought is a popular technique that significantly improves answers. Unfortunately, the latest version of GPT-4 did not generate intermediate steps and instead answered incorrectly with a simple "No."
Code generation has also gotten worse.
The team built a dataset with 50 easy problems from LeetCode and measured how many GPT-4 answers ran without any changes.
The March version succeeded in 52% of the problems, but this dropped to a pale 10% using the model from June.
Why is this happening?
We assume that OpenAI pushes changes continuously, but we don't know how the process works and how they evaluate whether the models are improving or regressing.
Rumors suggest they are using several smaller and specialized GPT-4 models that act similarly to a large model but are less expensive to run. When a user asks a question, the system decides which model to send the query to.
Cheaper and faster, but could this new approach be the problem behind the degradation in quality?
In my opinion, this is a red flag for anyone building applications that rely on GPT-4. Having the behavior of an LLM change over time is not acceptable.
Have you noticed any issues when using GPT-4 and ChatGPT lately? Do you think these problems are overblown?