MLB is cooked
If your baseball understanding is a spreadsheet, @TheJudge44 is your guy
If you actually watch baseball, you know @Mariners#CalRaleigh turned in a season no catcher has ever come close to
Judge was great. Cal was irreplaceable. "Valuable" is not "cell A1"
@rryssf Applying evolutionary forces to model refinement is a good idea
This work also has some parallels to how language and an internal monologue (or equivalent for people without one) could have translated into evolutionary advantages
The probability of the LLM making a bad decision for each action it takes is non-zero. Assuming programmers of equal skill, the difference is one or both of these:
A) Spending time to provide the LLM context, rules, and guidelines before using it
B) Carefully checking the LLM’s output to correct any bad decisions before they snowball
A and B are also non-binary, so one can do varying degrees of a good job on them to get different results
Some actual data-driven, practical advice for students using AI to get ready for AP test season, including specific recommendations for which models to use for which subjects. Edu AI benchmarks that mean something!
AP exam season is looming! 😰 Feeling the pressure? We tested how well popular AI models can help you get a 5. The results might surprise you 🧵👇
https://t.co/t4t81FQ2O6
#APTests#AI#EduTech
1/8
@athyuttamre Looks great! previous_response_id is killer. It always felt weird having to send so many strings back and forth. Same for hosted tools. It always felt weird that the model would need me to call basic tools for it.
Excited to try it out!
@tunguz If only some clever sci-fi author like Orwell or Stephenson or Asimov or Dick or Gibson or Bradbury or Vonnegut or Huxley or Vinge or Heinlein had warned us
It seems unlikely we'll see true automation of software engineering until AGI. It also seems unlikely AGI doesn't result in automating software engineering. So:
P(engineering_automated | not_AGI_yet) ~= 0
P(engineering_automated | AGI) ~= 1
So is your assessment P(AGI) in your lifetime is 25%?
In the meantime, what's most likely is dramatically higher engineering productivity for the developers who embrace AI tools. Which is what we're seeing.
Tbh I was surprised to see these results, given most LLMs did reasonably well creating multiple-choice questions. Something about creating a quiz really seems to throw them for a loop... digging into why soon.
Now live on https://t.co/hOvVeLGsuB: Quiz Composition benchmarks! Unlike MCQ Generation, frontier models struggle with most aspects of creating good quizzes. More details soon; for now, head over to https://t.co/hOvVeLGsuB and see for yourself. What stands out to you?
AI-generated MCQs can look good but still fail students. A key issue? Bad distractors—wrong answers that are too obvious to provide effective assessment. Measuring distractor quality is a critical tool for using AI effectively in the classroom
We're benchmarking AI for education. What does that mean? Let's dive into an example.
Suppose you're an AP World History teacher who wants to use an LLM to create a quiz question for your students... 🧵
1/7
One takeaway from our recent comparison of Claude 3.5 and 3.7 for edu:
Both Claude versions fail because they're optimized for capability demos, not educational outcomes. Simply being able to generate questions isn't enough. We need models that create content that actually helps students learn. #EdTech
Excited to announce a new project! In EduBench, we aim to comprehensively benchmark AI in educational settings.
We're starting with multiple-choice questions. Much more to come, including our own efforts to create best-in-class LLMs for edu applications. Follow for more!
🚨 AI in education is at a crossroads. Too many AI tools look helpful but actually mislead students, fail to align with curricula, and give a false sense of competence.
We’re launching EduBench to demand real educational impact from AI. 🧵👇
NEW POST
My colleague Bharani Subramaniam has started to write patterns from our recent work building production Gen AI applications. We begin with Evals - ways of assessing if they are working effectively.
https://t.co/NhEPQLMokX
Father of 2 here (and I couldn't imagine having more)
The right unit for thinking about parenting isn’t dollars. It’s hours
Child credits fall short because they don’t save time. Free child care does, because it's just there. Like a utility. You don't have to worry about it to use it
Most people don’t skip having kids due to a lack of money (income and fertility are negatively correlated). They’re time-constrained
Modern parenting means arranging summer camps, coordinating playdates, filling in for a declining education system, ensuring the “right” trajectory for college. Plus cooking, cleaning, laundry, chauffeuring, etc. Add in always-online work culture and two-income households (a net good, by the way)... it’s a lot
“Time is money,” sure. But it takes time to convert child credits to childcare. Parents are out of time
Pronatal policy should focus on giving parents time. Remove entire categories of worry. Don't just make having a kid "worth it" with money
Incidentally, AI’s promise for fertility isn’t extra productivity to fund retirees. It’s giving parents time: handling scheduling, homework, chores, even babysitting. It's making being a working (and sleeping) parent more possible in a 24-hour day
@rowancheung This shows 1 of 2 things re:“it’s autocomplete on steroids”:
- There’s more going on in these models than we can see beyond token prediction
- Predicting the next token is sufficient to create HL intelligence (this isn’t HL yet, but looks like a straightforward path to get there)
Strong agreement: Ed Dept's AI in education report hits the nail on the head. AI isn't just a cost-cutting tool; it's the key to unlocking personalized learning at scale. We can finally tailor education to each child's needs, not just the average student. https://t.co/zF6SEFzx7D