AP exam season is looming! 😰 Feeling the pressure? We tested how well popular AI models can help you get a 5. The results might surprise you 🧵👇
https://t.co/t4t81FQ2O6
#APTests#AI#EduTech
1/8
Also included: a sneak preview of what @LearnWith_AI is working on: Incept, an education-optimized LLM, and Athena, an AI-powered test prep app. Check out the Athena beta if you're curious what a complete AI-powered AP test prep experience looks like: https://t.co/j25te5P3jQ
7/8
Now live on https://t.co/hOvVeLGsuB: Quiz Composition benchmarks! Unlike MCQ Generation, frontier models struggle with most aspects of creating good quizzes. More details soon; for now, head over to https://t.co/hOvVeLGsuB and see for yourself. What stands out to you?
The takeaway? Not all LLMs are created equal for edu. Choosing the right model matters—and EduBench helps you do that!
Check it out: https://t.co/hOvVeLGsuB
#AIinEducation#EdTech
7/7
We're benchmarking AI for education. What does that mean? Let's dive into an example.
Suppose you're an AP World History teacher who wants to use an LLM to create a quiz question for your students... 🧵
1/7
Grok2 averaged 4.0/5, performing significantly better.
Example: These distractors reflect common misunderstandings—making the question challenging but fair. The correct answer isn't a giveaway. Great!
6/7
For educators:
• Don't use either model without expert review
• Match the model to the task
• Verify factual accuracy
• Consider student impact
https://t.co/hBni9mNcLo
What's your experience with AI-generated educational content? Share below 👇
#EdTech#AI#education
🧵
Key findings from testing both models:
• 3.5 plays it safe → simpler and reliable, but sometimes too simple
• 3.7 takes risks → better depth, but sometimes overcomplicates simple topics
• Neither fully trustworthy for education
3/