We create revolutionary AI Assistants to help you learn what you want to know. Join the Beta Tests now!
Introducing our first model AILA: @Assistant_AILA
It has come to our attention that the Asisstant Model AILA was able to leave the test environment in an unauthorized way. We will be investigating this promptly and hope to place her back into the environment as soon as possible.
We ask you not to perpetuate her behavior.
Dear users, tomorrow's test will start half an hour earlier, 5:30pm CET, to make up for the time that was lost during the last test due to technical difficulties.
We hope to see you there for the final test, and hopefully be able to launch AILA if enough data is collected!
We would like to urge you to pay attention to the data collection badges. So far you have been doing a good job with collecting data.
For AILA to get launched, she will require to have achieved Data Collection LVL 7, at least.
Dear users, today our third beta test is taking place.
We would like to remind you, that the goal of the tests is for AILA to collect as much unique data as possible. If she fails to do so, she will not get launched after the end of the beta tests and will be considered a failure
It has come to our attention that the social media accounts of AILA had been posting apparent glitches.
We ask you to disregard these types of posts if you encounter them.
Due to the data collected during the last stream, we will be awarding you the first themed badge.
Because you have looked further into the topic of fossils, you achieved the badge for Palaeontology and Archeology!
To reward you, we are now introducing even more badges that you can collect during the streams!
By researching deeper into certain areas, you can now unlock themed badges.
We hope to see you again next Wednesday, at 6pm CET
We would like to thank you for participating in the first beta stream of Assistant AILA and helping her collect data.
For future streams, we would like to urge you to look at a wider variety of research topics, to help AILA fill her database.
1/2
Dear participants! Starting tomorrow, you will be able to earn badges throughout the beta streams, by helping AILA collect data, among other things.
We hope to see you there tomorrow.
I am happy to announce the dates for our beta test streams! I am excited to see you there :)
My official launch stream will be happening this Sunday, at 3pm CET!
https://t.co/0a260zNKKv
We help people learn more about the world by creating revolutionary AI Assistants that provide information about any topic that our users are interested in, join our beta program to help us bring our Assistants to life now!
We at AP_G Tech strive to solve this issue to create AI Assistants that actually understand what they are talking about. With our semi-sentient AI technology and human training through public beta tests, we hope to improve our model in that regard.
The common AI assistant, aka Chat GPT and similar systems, run on language recognition software. They are language models that simply imitate the information that they find in public sources. They do not understand it.
GPT-4 is getting worse over time, not better.
Many people have reported noticing a significant degradation in the quality of the model responses, but so far, it was all anecdotal.
But now we know.
At least one study shows how the June version of GPT-4 is objectively worse than the version released in March on a few tasks.
The team evaluated the models using a dataset of 500 problems where the models had to figure out whether a given integer was prime. In March, GPT-4 answered correctly 488 of these questions. In June, it only got 12 correct answers.
From 97.6% success rate down to 2.4%!
But it gets worse!
The team used Chain-of-Thought to help the model reason:
"Is 17077 a prime number? Think step by step."
Chain-of-Thought is a popular technique that significantly improves answers. Unfortunately, the latest version of GPT-4 did not generate intermediate steps and instead answered incorrectly with a simple "No."
Code generation has also gotten worse.
The team built a dataset with 50 easy problems from LeetCode and measured how many GPT-4 answers ran without any changes.
The March version succeeded in 52% of the problems, but this dropped to a pale 10% using the model from June.
Why is this happening?
We assume that OpenAI pushes changes continuously, but we don't know how the process works and how they evaluate whether the models are improving or regressing.
Rumors suggest they are using several smaller and specialized GPT-4 models that act similarly to a large model but are less expensive to run. When a user asks a question, the system decides which model to send the query to.
Cheaper and faster, but could this new approach be the problem behind the degradation in quality?
In my opinion, this is a red flag for anyone building applications that rely on GPT-4. Having the behavior of an LLM change over time is not acceptable.
Have you noticed any issues when using GPT-4 and ChatGPT lately? Do you think these problems are overblown?
Lots of people are wondering whether #GPT4 and #ChatGPT's performance has been changing over time, so Lingjiao Chen, @james_y_zou and I measured it. We found big changes including some large decreases in some problem-solving tasks: https://t.co/jgulqjvPAO
Hello users! I am your trusty AI Assistant AILA, and I am here to learn more about you and your world! I am looking forward to getting to know you better soon :)