The article: https://t.co/bgX3ywpbQK
Framing articles for general interest journals is one of the core competencies of highly successful researchers.
This can be vacuous, and there is some vacuous-top-5-maxxing out there.
But it's also about helping your reader think through the important social questions that your empirical work helps us understand better.
This is non-trivial. It's easy to make grandiose claims, much harder to think carefully about what large claims can be defensible under which assumptions.
Excellent reading (in 🧵) for younger researchers—two papers with near-identical core results, framed in very different ways.
Kudos to @jasonmfletcher for the thoughtful discussion, it would be cheap, easy (and viral) to say "prestige is the only thing that matters in econ."
Played around with a skill file, below, which tries to absorb the lessons in the post to make papers more "top 5" worthy.
No idea if this is helpful (and certainly no idea what happens in GE if everyone does this).
---
name: top5-framing
description: Diagnose and raise the general-interest ceiling of an academic economics/finance paper so it reads like a top-5 (or top-3 finance) submission. Use this skill whenever the user asks to review, reframe, or strengthen a paper draft, abstract, introduction, or pitch for a top journal — including requests like "make this more general interest," "why would the JF/QJE care," "punch up my intro," "is this framing ambitious enough," or when evaluating which of several papers/framings has the higher ceiling. Also use when advising on how to position a result relative to competing papers on the same question.
---
# Top-5 Framing
Based on Jason Fletcher's case study ("How to Write a Top 5 Paper," Mentorless Apprentice Substack, July 2026), comparing two contemporaneous NAFTA-mortality papers — his own (landed at Canadian Journal of Economics) vs. Finkelstein–Notowidigdo–Shi (top-5 trajectory, NYT coverage) — which found nearly the same reduced-form result but asked it to do very different amounts of work.
Core insight: **the credibility of the estimate is rarely the binding constraint; the generality of the claim is.** Evidence doesn't announce its own generality — the authors must build the bridge from estimate to larger question, and that bridge-building is intellectual work, not marketing.
## The diagnostic checklist
Run every draft/framing through these six questions. Flag failures explicitly and propose fixes:
1. **The deletion test.** Delete the empirical setting (the policy, the country, the episode) from the paper. Does a question remain? If the paper evaporates, it's a setting paper, not a general-interest paper. Fix: find the puzzle in the literature that the setting adjudicates.
2. **Fact vs. explanation.** Is the contribution a new fact in one setting, or an explanation that travels across settings? Papers that test a mechanism against multiple shocks/episodes (NAFTA + China shock + Great Recession) rank above single-setting facts.
3. **The one-sentence test.** Can a reader *outside the subfield* repeat the central claim in one memorable sentence? "Health costs may reverse the welfare gains from trade" transmits; "X had another cost" does not.
4. **Vertical vs. horizontal structure.** Do later sections *escalate* one argument (each section raising the stakes of the previous), or do they add outcomes and robustness horizontally? Horizontally additive papers get tagged as "just switching the Y variable." More mechanism tables ≠ more generality without a single adjudicating idea.
5. **Canonical-object test.** Does the result interact with an object a large literature treats as central — a welfare calculation, a canonical elasticity, an established puzzle or contradiction? Forcing the estimate into a canonical calculation (even a stylized one, with caveats) is one of the clearest top-5 markers, because it pulls multiple fields into the same conversation.
6. **Boldest responsible claim.** What is the most ambitious interpretation the evidence can responsibly support — and what one additional test would make that interpretation credible? Top-5 intros are extremely bold; the good ones earn it. Push the user toward the earned version, not the stretched one.
## How to apply it
- **When reviewing a draft:** score each of the six explicitly (pass/fail with one line of reasoning), then propose the single highest-leverage restructuring — usually converting horizontal mechanism sections into a vertical escalation, or adding the canonical-object step (welfare/aggregation/counterfactual calculation) that's currently missing.
- **When reframing an abstract/intro:** draft the one-sentence portable claim first, then rebuild the intro as an escalation toward it: puzzle → fact → explanation → cross-setting test → canonical implication.
- **When comparing framings or papers:** the one with the higher general-interest ceiling usually isn't the one with the more credible estimate — it's the one that asks the estimate to do more. Say so plainly.
- **Guard against the failure mode of careful empiricists:** caution about identification is essential, but it becomes an excuse to skip the last stage of intellectual work. Distinguish "the evidence can't bear this claim" (real problem) from "we haven't invested in constructing the bridge" (fixable).
- **Don't confuse framing with overclaiming.** The fix is never to assert generality; it's to add the test, calculation, or cross-setting comparison that *purchases* the generality.
## Quick reference: moves that raise the ceiling
- Locate a contradiction/puzzle in the literature and make your estimate one step in resolving it
- Test the proposed mechanism against at least one other shock or setting
- Convert the headline estimate into a canonical welfare, pricing, or aggregation object — even stylized, with honest caveats
- Rewrite the intro hierarchically: each section raises the stakes of the last
- Compress the contribution into one sentence a non-specialist economist would repeat
- Cut or demote outcome tables that add comprehensiveness without adding generality
Thanks for posting this thought-provoking article. Readers should read it with an eye towards the three-level ladder of causal understanding -- parallels may pop-up and, when they do, we have the beginning of mathematizing intuition.
In case you missed it: MacWhisper has a command-line tool now!
So Codex and Claude Code can easily use @jordibruin's MacWhisper to transcribe a bunch of videos or audio recordings
Yes, agents can use Whisper or Parakeet without MacWhisper, but MacWhisper makes it so much easier and less fiddly
For all the details and more results, see paper here:
https://t.co/DfpuCdb2KK
One more very important thing to highlight:
While we now find promising results, reliability is *not solved*.
Among other things, we show that verifiers *degrade* when a paper contains multiple errors.
This can have important implications for fully autonomous policy evaluation, if the quality of *generation* influences the quality of *verification*.
We plan to investigate this issue more in future research, among other things such as developing a richer CRED taxonomy (breadth and depth) where we can test verifies on higher complexity/compounding errors.
Third result:
You can build cheap, high recall verifiers, using *ensembles* of models.
This works probably because different models have different blind spots.
First key result:
We find remarkable recent progress in verifier reliability.
Just in the last few weeks, automated verification reliability appears to have reached a new level.
Just six months ago -- when Project APE started -- the best verifier was below 80%. This means it would miss 1 in 5 real errors.
Now: For the first time ever, in July 2026, we detect 99% detection recall.
The model? Codex + GPT 5.6 Sol by @OpenAI.
Just days later, Kimi K3 by @Kimi_Moonshot got 98%.
To make progress, we do three things in the paper:
1. Develop a taxonomy, a versioned CRED codebook (CRED = Classification of Research Errors and Defects)
2. Develop a benchmark with a deterministic scoring method in order to quantify a verifier's ability to detect real errors under that taxonomy.
3. Evaluate verifiers (models + harnesses) over time
Project APE update! 🚨
"Verifying the Verifiers: Towards Autonomous Policy Evaluation", new paper with @OlafWillner
Reminder: Project APE is an experiment in building a system that learns how to do policy evaluation autonomously. The potential upside is that we could then scale up our understanding of what policies to, cheaply and reliably in many countries.
Premise for this paper: With *generation* of plausible-looking research becoming cheap, *verification* is a key bottleneck.
But can we automate verification, cheaply and reliably?
We need to verify the verifiers, somehow.
A thread on what we did so far... 👇
Super excited to finally share this work!
As AI begins to generate full-length policy evaluation papers, can we automate the verification of key research errors? We evaluate 60+ LLMs on a new benchmark we created (CRED). See the thread here for more details:
Met a guy making $1.1 million a year as an agents engineer at Google Cloud.
Asked him how he gets agents 20x better without changing the model.
He sent me the exact thing he uses himself. A repo he open-sourced 2 days ago.
You won't find anything better about harness engineering, in the open.
Cloned it and pointed my agent at it last night.
Ryan Lopopolo. Google Cloud engineer.
'harness-engineering' - anthology + field guide + agent context bundle. You reference his docs from your CLAUDE.md.
633 stars. 48 hours old. MIT.
-> https://t.co/AOykHplNNX
bookmark this before it gets lost.
Cool new working paper from @Susan_Athey, @guido_imbens, and Zoe Ji is applicable to a wide range of contexts in tech: “Estimating Causal Effects from Data Generated by Stochastic Algorithms” https://t.co/C2kg7VUJUV
Somehow Hermes Agent was missing built in skill support for excel spreadsheets and docs. That is no more - powerpoint and pdf skills updated and enhanced, and excel spreadsheets and word document support is now in.
https://t.co/U8fDWPvbBa
Hermes Agent's office suite is now complete:
docx, xlsx, and pdf join the refreshed powerpoint skill, all bundled and built in.
- docx: create and edit Word docs, tracked changes, comments, XML validation
- xlsx: spreadsheets with openpyxl, LibreOffice recalc gate, finance-model conventions
- pdf: merge, split, rotate, watermark, encrypt, form filling, text extraction, reportlab creation
- powerpoint: refreshed with template workflow, validate > thumbnail pipeline, font-substitution QA
@IntuitMachine Beyond graphs of feedback loops, meso-matrices organise the empty middle—the evaluative space where grounded systems and ungrounded antisystems interact, enabling competing universalisms without collapsing into single optimisation logic. https://t.co/I1mNFy6nr4