Our new PLOS One paper gave ChatGPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6 the same 20 randomized trials and the same CONSORT checklist. They disagreed substantially, with mean compliance scores ranging from 54.7% to 80.8%. The AI model a journal picks can change how a manuscript is judged, so standardized protocols and validation against expert human review need to come before AI-assisted compliance screening. https://t.co/iCdELHBkzZ
Interesting study. Thanks for sharing. Crude 4-year mortality, the primary outcome, was 0.43% in vaccinated vs. 0.55% in unvaccinated. That's an absolute difference of about 0.12%, or an NNT of roughly 845. The authors present this as a safety finding (no increased long-term mortality) and note likely residual confounding from the healthy-vaccinee effect. It's reassuring on safety, but 'massive life saving benefits' goes well beyond what the authors claimed.
The 29% reduction in the 6 months after vaccination comes from a self-controlled case series of deaths only, so it cannot be used to calculate an absolute benefit. It also showed a similar drop in deaths from accidents, suicides, and other injuries, which points to a selection rather than a protection effect.
The MUTTON-HF food-is-medicine trial (p=0.025) has a resampling fragility of 0.38 and a GFQ of 0.0097. An estimated 38% of replications would lose significance, and just 2 outcome changes would reverse it. I recommend cautious interpretation. Here's my full analysis: https://t.co/Op9yHmLstv
Extending neutrality boundary robustness to meta-analysis. Across 161 trials, median nb = 0.147, and p-values explained only 12% of the variation in nb. A 0–1 score for how far evidence sits from no effect. https://t.co/0E10TvKt4g
64 best-selling multivitamins, 25 nutrients tracked: products contained anywhere from 7 to 24 of them, vitamin B12 ranged from 100% to 33,333% of the Daily Value, and 47% megadosed at least one nutrient. "Multivitamin use" is not one exposure. Research studies that only state "multivitamin use" provide no meaningful information on what was actually consumed.
https://t.co/KPWsQRLjDr
The significance classification in OFACAR is unstable under all four forms of statistical fragility.
1. Analysis fragility. For the primary-outcome table [[120, 39], [136, 25]], Pearson χ² p = 0.0442 (significant) and two-sided Fisher’s exact p = 0.0507 (nonsignificant). The significance classification depends on the test.
2. Resampling fragility. Under independent binomials at the observed rates and arm sizes, with two-sided Fisher as the rule, P(p < 0.05) = 0.499368. A new sample of the same size from that same population is as likely to be called significant as not. If we drew 100 new samples of the same 320-patient size from that population, about half would be labeled significant and half would not.
3. Perturbation fragility. The bidirectional FI is 1: one within-arm outcome toggle changes two cells and flips the classification. More importantly, the fragility distance (FD) is also 1, but it is an L1 distance, not an index: adding or removing one patient from a single cell is enough. Observed Fisher p is 0.0507, so these edits flip a nonsignificant table to significant. Adding one OFA non-event gives [[120, 40], [136, 25]], N = 321, Fisher p = 0.0379. Removing one control non-event gives [[120, 39], [136, 24]], N = 319, Fisher p = 0.0355. Stopping at N = 319 or continuing to N = 321 could have changed the Fisher classification.
4. Scaling fragility. Hold the observed rates fixed and use Pearson χ², the test that makes the baseline table significant. The sample-size fragility multiplier (SFM) = 1.054328. Shrinking N by that factor reaches the p = 0.05 boundary at [[113.816556, 36.990381], [128.992097, 23.711783]], N = 303.51, χ² = 3.84146, p = 0.0500. The nearby integer table [[114, 37], [129, 24]], N = 304, has Pearson p = 0.0550 (nonsignificant). With the same outcome ratio in each arm, a 5.15% reduction in sample size flips the Pearson classification.
Robustness is the distance from therapeutic neutrality (relative risk = 1). RQ = |ad − bc| / (N²/4) = 2304 / 25600 = 0.0900, an intermediate distance from no effect. That does not support a strong effect-exists claim.
Overall, complete statistical evidence, the p–fr–nb triplet, shows: p (significant by the researchers' analysis but not under Fisher); fr is unstable; nb is intermediate. With unstable classification on all four fragility forms, the statistical evidence is inconclusive. Changes in clinical practice generally should not rely on these findings. Readers should remain very skeptical.
RQ can be obtained at https://t.co/YDQyxvmw5L. FD, SFM, and resampling fragility were computed from the definitions in the documentation there, not from the on-site calculator.
Clinicians consult AI about real patients every day, but there's no standard way to report what happens in a single consultation. Here's my take on the AI consultation report, for documenting when AI helps or fails at the bedside.
https://t.co/E3YWVmg70u
Druckenmiller: There’s a reason I moved from an English major to being an economics major. I’m not embarrassed by it. I write everything using AI now for the same reason I use a calculator when I do math problems. I don’t know why this is relevant. My name is on the piece. It’s my message.”
@brandonbaumet I’ve seen very similar responses on rare occasions. Worst was after I cut-n-pasted a prompt off X that basically told Claude to play the role of a no BS advisor.
Global Fragility Index explained in the context of the Complete Evidence Standard for medical research:
the p-fr-nb triplet for complete statistical evidence.