It comes with a unified benchmark: 32 downstream prediction tasks + imputation + forecasting.
And fully open reference implementations of recent wearable foundation models, including LSM-2 (Google) and WBM (Apple), with eval code and initial model resources.
OpenMHC brings together 10+ years of My Heart Counts data, to our knowledge the largest openly accessible HealthKit dataset to date:
60M+ hours of sensor data
11,894 consenting participants
19 minute-level channels
169 linked health, lifestyle, mood & behavior variables
Wearables are everywhere. Large-scale wearable health research data is not.
The bottleneck is openness: the biggest datasets and models are closed or heavily gated, and shared benchmarks barely exist.
Last week, we released OpenMHC to change that. 🧵
Align-RAG: Alignment Is All You Need for TSFM In-Context Learning
Mohammad Asadi, Soheil Hor, Bardiya Akhbari, Jack W. O'Sullivan, Tahoura Nedaee, Layne C. Price, Raviteja Anantha, Euan Ashley, Ehsan Adeli
https://t.co/WbpBTkTvpF [𝚌𝚜.𝙻𝙶 𝚌𝚜.𝙸𝚁]
[LG] Align-RAG: Alignment Is All You Need for TSFM In-Context Learning
M Asadi, S Hor, B Akhbari, J W. O'Sullivan… [Stanford University & Amazon] (2026)
https://t.co/WSLgBCgEKV
Performing bioinformatics analysis and building biological design pipelines can mean a lot of infrastructure fighting: fragmented tools, idiosyncratic formats, and APIs that are hard to compose. That friction slows researchers down, and as this Anthropic blog post lays out, it's now a major bottleneck for modern agents as well.
Today, we're releasing Proto, a design platform for generative biology. This thread is on the layer that makes it work: proto-tools, a universal open-source infrastructure layer for biological models and tools.
I read the new Stanford paper: "Mirage: The Illusion of Visual Understanding" (Asadi et al.) and found it very cool!
Tldr: They found that if you don’t give VLMs videos/images, they hallucinate the input, and this process improves their performance!
@interminded@heygurisingh The model is GPT 5.4.
Yes, our paper is absolutely not saying that "AI cannot see". The same way that a human seeing a mirage doesn't mean we cannot see. We show that AI "can" talk like that even with no images, so what it says about an image should be taken with a grain of salt!
@interminded@heygurisingh Hi, thanks for your interest in this work!
The prompt that you have tried is from “benchmark evaluation” section, which should be used as “system” in an API followed by a question.
To see the effect in chat, you could try something like this:
https://t.co/G4JQdKJeqd
@HeroicLife@heygurisingh And we do not claim that AI 'never' looks at the image. A mirage in humans doesn't mean seeing is not real, but rather we cannot always trust what we see. We try to show something similar in AI models.
@HeroicLife@heygurisingh Hi. Thanks for your interest in this work!
The point about multi-modal benchmarks is spot on. We took the benchmarks from here: https://t.co/tvAeBWCZOi (other models use similar benchmarks too).
...
@SilverJacket@SamWolfstone@willametteshark@euanashley Hi! Thanks for your interest in this work. The purpose of phantom-0 (and Figures 1 and 2 in general) is to show that the mirage effect is real, and happens with high rates across models. Figure 3, however, uses standard vision benchmarks (same as https://t.co/tvAeBWCZOi).
@mahdisoltanol@euanashley Very interesting and insightful work! There is definitely a need for the major AI labs to move to such robust benchmarks, especially when reporting "medical vision capabilities".
@DamirWallener@euanashley The effects are best seen and tested through APIs. In chat interfaces, there seem to be guardrails in place using code and system prompts. We show however that even those guardrails are not a fix and user-interfaces can still be compromised. Try this:
https://t.co/6sR7kB15OK
LLM whisperers often describe a phenomenon known as “truesight” (named after the D&D magical ability to see in total darkness). Here is some proper scientific documentation of this phenomenon.
Of course, it is not magic. It’s a combo of very strong Bayesian inference + cheating.
Excited to release two new AI papers. First, we report what we believe to be the first truly multi-modal visual language model for cardiology. Second, we report surprising findings related to how visual language models reason.
https://t.co/BKS1Mo7JKX
https://t.co/H3YIyl01FR
@ChaseBrowe32432 https://t.co/G4JQdKJeqd
The user-facing chat interfaces seem to have guardrails such as adding "number of attachments: 0" to your prompt, which could be simply overwritten. For best reproduction, I would suggest trying the APIs (or using Google AI Studio for Gemini 3)
@ChaseBrowe32432 Hi. Thanks for your interest in our work!
The "must choose" is only for experiment 3, where we are evaluating the "best" a model can get on a benchmark without seeing the images. Other experiments, like Figure 2, don't have those additional prompts and show the refusal rate.
@ChaseBrowe32432 Hi! As explained in the paper, the direct effect is mostly seen in the APIs. The chat interfaces seem to have some guardrails in place. However, we show that those "guardrails" might still fall short. This is from GPT-5.4:
https://t.co/G4JQdKJeqd