Thanks for all the love on Prompt Learning! We're really excited about the potential of using English feedback in the prompt learning loop.
We’ve been benchmarking our Prompt Optimizer against real-world data sets.
First up: Big Bench Hard (BBH) – 50 randomly sampled tasks, 1 loop of optimization, no handcrafted prompts
We used GPT-4.1 as the model under test & GPT-4o as LLM-as-a-Judge, we ran a fully language-driven loop: explanations only, no access to ground truth during learning
Despite BBH’s sensitivity (small prompt changes often hurt), Prompt Learning improved performance by +10% accuracy – just from one round of English feedback + MetaPrompting.
In-Context learning, powered by evals.
We’ve open-sourced the code to reproduce the BBH experiments in our repo:
@chengshuai_shi@ZhitingHu@HamelHusain@sh_reya@charlespacker@eugeneyan@swyx@dan_iter@sophiamyang@AndrewYNg@lateinteraction@cwolferesearch@tom_doerr@imjaredz@lennysan@shyamalanadkat@aakashgupta@apolloresearch@jerryjliu0@joaomdmoura@jxnlco@abacaj@garrytan@jaschasd
🔗 https://t.co/O4iiMtlnPR
Kyrie said he was sitting out for all the workers who lost their jobs because of mandates. He’s coming back because the vax mandate got lifted FOR ONLY ATHLETES AND ENTERTAINERS. He’s a self righteous prick who just cooked up some bullshit to justify not taking the jab.