PhantomEnvironments is coming out soon! I'll be presenting in SF at @SnorkelAI Frontier Data Summit (DM me for an invite code) and @COLM_conf workshops! I have more announcements for PhantomEnvironments, but gotta do ICLR first.
Synthetic RL envs to train off-the-shelf LLMs as multi-turn search agents. These environments are synthetic to the extreme and "air-gapped". Absolutely no distillation, train-test knowledge leakage, or benchmaxxing. Yet the benchmarks get maxxed. We were able to separate learning "agentic search skill" from memorizing knowledge... as if they are complementary axes for agents.
@ShcChy@iclr_conf@NeurIPSConf Small part of the larger discussion, basically just reiterated what the blog said, highlighted that it came from the blog, and shared the link to the source material… dont think thats pretty misinformative but okay 🙄
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE
@TmlrOrg Interesting process… I was just speculating on what review processes may look like with the deluge of submissions in the AI era… posted this in relation to ICLR:
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE
We think it’s great that lots of people are reviewing and scrutinizing DeepSWE. It being our first benchmark, we wanted to invite public collaboration and discourse, which is why we shared all our tasks and grading results in the open.
Since then, we've made lots of improvements in our methods internally and I'm very excited about the problems we’re better able to solve as a result. Notably, we’re at a point again where no benchmarks accurately reflect the capabilities of current models, and our upcoming work will be addressing this.
Seeing a lot of the benchmark folks respond to this… some say the findings resonate with real problems worth addressing… some thankful for the audit findings… and some posing real questions about who audits the auditors… Im going to pull some responses and start a thread:
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
Seeing numbers of ~50k submissions for @iclr_conf. That's insane!
Yes, probably a lot of AI papers. And yes, it'll probably need AI review at that scale. But @NeurIPSConf just ran into this exact wall, and the results were... interesting.
Based on some blogs that have come out, NeurIPS screened all 969 submissions through an AI detector. Out of those, 42.7% of submissions initially scored in the 90-100% AI-generated range, and 273 papers hit a full 100%. The chairs re-ran it with a narrower text window and that number dropped to 12.7%. Same papers. Same tool. One settings change, and the flag rate fell by 70%.
Resultantly, 178 papers were desk-rejected with no appeal and 123 were told to produce version histories or also be desk-rejected. Then a rejected author ran the track chairs' own recent papers through the same detector and got back ranges from 24% to 69% of AI-generated content.
That's with 969 submissions.
Now let's consider ICLR. There are reports that people's submission numbers are getting real close to 50k. The one that I saw plainly stated was 47647.
Compared to last year, there were 19,525 valid submissions to ICLR, with 779 desk rejected and 5,042 withdrawn. This left 13,763 papers that needed a decision.
To do this, ICLR organized 76,139 reviews from 18,054 reviewers. They also ran their hallucinated-reference checks through at least three humans per flagged paper before any desk rejection, putting in a check to any AI review processes.
Now what is this going to look like with a potential 2.5x jump in volume?
At ~50k submissions, ICLR is definitely going to have to use some AI review processes, but can the human-review layer scale as well, or is it something that may end up getting dropped?
This is speculation at this point, and is based solely on the abstract submissions so far, but the final paper deadline is Sept 25th, and it will sure be interesting how this all plays out.
If you want some advice though, maybe save your drafts, keep your version history, and be specific about any AI use in your attestation. That could come in handy if you get a rejection...
Anybody have thoughts here? Gonna be an impressive feat nonetheless. 🤔
NeurIPS Info: https://t.co/heCi0dAZvU
ICLR Info: https://t.co/eXY5mwUhUA
ICLR Conference Page: https://t.co/ajHoQG6dTY
Graph: https://t.co/Q1eVrsIIJE