GPT-4 is getting worse over time, not better.
Many people have reported noticing a significant degradation in the quality of the model responses, but so far, it was all anecdotal.
But now we know.
At least one study shows how the June version of GPT-4 is objectively worse than the version released in March on a few tasks.
The team evaluated the models using a dataset of 500 problems where the models had to figure out whether a given integer was prime. In March, GPT-4 answered correctly 488 of these questions. In June, it only got 12 correct answers.
From 97.6% success rate down to 2.4%!
But it gets worse!
The team used Chain-of-Thought to help the model reason:
"Is 17077 a prime number? Think step by step."
Chain-of-Thought is a popular technique that significantly improves answers. Unfortunately, the latest version of GPT-4 did not generate intermediate steps and instead answered incorrectly with a simple "No."
Code generation has also gotten worse.
The team built a dataset with 50 easy problems from LeetCode and measured how many GPT-4 answers ran without any changes.
The March version succeeded in 52% of the problems, but this dropped to a pale 10% using the model from June.
Why is this happening?
We assume that OpenAI pushes changes continuously, but we don't know how the process works and how they evaluate whether the models are improving or regressing.
Rumors suggest they are using several smaller and specialized GPT-4 models that act similarly to a large model but are less expensive to run. When a user asks a question, the system decides which model to send the query to.
Cheaper and faster, but could this new approach be the problem behind the degradation in quality?
In my opinion, this is a red flag for anyone building applications that rely on GPT-4. Having the behavior of an LLM change over time is not acceptable.
Have you noticed any issues when using GPT-4 and ChatGPT lately? Do you think these problems are overblown?
J'ai vu des ingés porter des projets à + de 800K alors qu'eux en touchaient 2K5 pour 50h/semaine, être contents de choper 3% d'augment annuelle, alors que la boite claquait des 10K dans des formations pour qu'un "coach" vienne te dire si t'es un collaborateur sanglier ou chaton
📢 J'ai l'honneur de vous présenter le renouveau de l'IA, grâce au travail acharné de mon frère, ce génie, voici CHATCGT, votre IA marxiste !
Elle est vener, elle aime pas Macron, c'est CHATCGT
📢☭ https://t.co/ZLAnHzh3rr ☭
@ekit0Kun@U_K_Shariban Oui j'ai quelques 1CC sur ma chaîne Youtube dont je me sers pour attester de mes scores sur https://t.co/kEpddCb8QP. S'il y en a un qui peut t'intéresser dit moi. https://t.co/l2PW3FBv2a
Récemment j’ai fait un tweet sur ma non envie de bosser avec des teams qui pratiquent scrum. Alors que je pensais que comme d’habitude ce tweet passerait inaperçu il a explosé (à l’échelle de mon compte) et ouvert un débat sur les qualités et défauts de scrum. #Thread
🔥 New Post: Announcing InAppBrowser - see what JavaScript commands get injected through an in-app browser
👀 TikTok, when opening any website in their app, injects tracking code that can monitor all keystrokes, including passwords, and all taps.
https://t.co/TxN1ezZX71
@AmazonFrance prime passe de 49 a 69 euros au prochain renouvellement. C'est trop cher pour financer un service de vod que je n'utilise pas. Pourquoi pas un abo séparé pour l'e-commerce ?
@_smontlouis Salut Steph, si tu es à la recherche d'un chat, il y a aussi le site de la SPA qui est très bien fait. Je pense que tu as sûrement déjà regardé mais j'aime bien le rappeler un peu partout quand je vois des projets d'achats. C'est ma petite contribution au monde :p
@taylorotwell Nice, thanks for the refresh ! I have an improvement idea regarding the scheduler section. I would love to be able to edit a job without having to re-create a new one. When I upgraded to PHP 8.1, I had to do a delete/create for each job with a higher risk of copy/paste mistake ^^