Is the "better writing" in the room with us?
Fable 5 wrote two novels, and we had readers find out
The good ✅: coherent 100K+ word plots and readers didn’t immediately stop
The bad 🚩: pacing, voice, and self-editing still break (+ see who swore at Claude 167 times 🤬)
👇🧵
ACL Sustainable Reviewing Policy: We are introducing changes in the @ReviewAcl reviewing and submissions. The changes will cap authors and introduce changes in the reviewing to keep our community sustainable. #NLProc
personally i am not really impressed when a computer is good at math — since a computer is math . i am a lot more intrigued by things like , for example , a rat being an exceptional cook — since that’s not really expected
I've written a blog post about our new method Progressive Point Matching, a simple approach to dense credit assignment for LLM RL. We improve training efficiency exponentially over standard GRPO on long-horizon tasks!
https://t.co/Z2mh8MW2ex
I asked GPT-6 Astra to mine a diamond in Minecraft using computer use, then went to sleep.
Woke up to a diamond in its inventory 🤯🤯🤯
Setup in the replies.
I'm moving to Nanyang Technological University in Singapore to start a new lab! We'll still be dedicated to adversarial QA, human-computer collaboration, probabilistic modeling, and the other fun hijinx my students and I have been up to, but now (hopefully)
bigger and better.
I'm agressively recruiting students and postdocs for 2027. I've tried to outline everything on my openings page, including a university-level postdoc with a deadline in October (but info sessions next week):
https://t.co/KV24BJzaRF
@hole_ghost@ManifoldMarkets nah its allowed, if you want there's an option to turn off market creator from betting. If you have a specific question in mind tho, I can create it for you
Wondering about the current state of CoT monitorability evaluation (a sticky research problem)? Check our our recent blogpost, where we break this down in detail 👇
https://t.co/0DU5CFBiu7
A study of private grant proposal submissions and publicly released awards from NSF and NIH finds a sharp rise in LLM use since 2023. At NIH—but not at NSF—proposals created with LLMs are more likely to be funded and to yield more publications. In PNAS: https://t.co/ONAA3DoFPW
@DimitrisPapail What did you find 5.6 Sol was good at for creative writing tasks?
We found it was better than Fable at giving feedback, but its prose was borderline unreadable/sloppish for long word counts.
https://t.co/vmOJBFraQH
Is the "better writing" in the room with us?
Fable 5 wrote two novels, and we had readers find out
The good ✅: coherent 100K+ word plots and readers didn’t immediately stop
The bad 🚩: pacing, voice, and self-editing still break (+ see who swore at Claude 167 times 🤬)
👇🧵
We used Fable to generate “love you dipshit”, a 130K word novel narrated by a profane AI named VERVE that's trying to survive its impending deletion.
Unsurprisingly, Fable loves its own book, claiming that it contains “satire at the level of the best DeLillo" and that “Eggers’s The Circle is a pamphlet next to this”.
Human readers are decidedly more mixed: some chapters are surprisingly funny, well-written, and emotionally engaging… but others drag due to slop, pacing, and VERVE’s overbearing snarkiness. Better than any other AI book we've seen (and we've generated/read millions of tokens of AI bookslop), but still far from supplanting a human author.
Check it out on our autofiction platform and let us know at what chapter you DNFed! We also have a slower-paced Fable romance, “Eight Tuesdays”, in case that’s more your speed.
AutoFiction is a research platform built in collaboration with Chau Minh Pham @chautmpham, Yapei Chang @YapeiChang, and Mohit Iyyer @MohitIyyer at the University of Maryland, College Park @ClipUmd@umdcs
Website: https://t.co/B8D1gfKm5C
GitHub: https://t.co/OUZXd2A2YK
Is the "better writing" in the room with us?
Fable 5 wrote two novels, and we had readers find out
The good ✅: coherent 100K+ word plots and readers didn’t immediately stop
The bad 🚩: pacing, voice, and self-editing still break (+ see who swore at Claude 167 times 🤬)
👇🧵
As with all books on our platform, both books are free to read, and our blogs provide additional information on harness design and the types of human feedback we used. We release:
🧾 traces
📈 timelines
🗂️ slop casebook
We need readers. Pick a book up and tell us about it!