Interested in program synthesis for creating random DNNs? and its application on automated testing?
Check our new work: “NeuRI: Diversifying DNN Generation via Inductive Rule Inference” with a Distinguished Paper Award @FSEconf!
w/ Jinjun, @YuyaoStarling, @LingmingZhang
In the past 6-mon release of HumanEval+ we have been improving its toolchain usability and dataset quality from v0.1.0 to v0.1.7 releases.
🔥 Now we release MBPP+, a new benchmark in EvalPlus v0.2.0: https://t.co/2vUM9b38s7
🧵
Introducing the EvalPlus leaderboard! https://t.co/H0aOtkzbnc
🔥28 models have been evaluated on coding HumanEval & HumanEval+
🔥7B CodeLlama outperforms ~16B models e.g. StarCoder&CodeGen
🔥Phind-CodeLlama-34B-v2 and WizardCoder-Python-34B-V1 as open models both beat ChatGPT
🧵
New Arxiv: https://t.co/i27lVI8dPD
GPT-4/PaLM-2 have both shown almost perfect performance on existing grade school math dataset. What about more challenging STEM questions, especially the ones which require specific theorems, like Stoke's theorem, Wiener Process, etc?
We welcome everyone to try out 📚𝐇𝐮𝐦𝐚𝐧𝐄𝐯𝐚𝐥+! A dataset to reflect the "real" correctness of LLM-generated code.
Using📚𝐇𝐮𝐦𝐚𝐧𝐄𝐯𝐚𝐥+ is the same as HumanEval. You can easily pip install it and evaluate in our prepared sandbox (optional). https://t.co/96YcDvNJpm
🚨 Evaluating LLM-generated code on datasets with just "3 test-cases" is NOT enough! 🚨
We built ✨HumanEval+✨: improving HumanEval with up to thousands of new tests to fully evaluate functional correctness of LLM generated code!
@JiaweiLiu_@YuyaoStarling@LingmingZhang
Strong PhD candidate in ML/AI:
“I have published four NeurIPS papers, two are first-author, one of which was a spotlight”.
Strong PhD candidate in PL:
“I have solved all exercises from the first two volumes of Software Foundations”.
I tell new PhD students to pick a research topic according to three criteria: (1) the problem should be important, (2) it should have a reasonable chance of being solvable, and (3) you should personally have a unique edge.