Friday, August 1, 2025, 5:36 PM, San Francisco time.
An email from Y Combinator arrived. They wanted us in San Francisco for an in-person interview the following Monday.
Three days later, Marc and I were sitting in the YC offices, talking about Kosmico.
We didn’t get in. It was disappointing, but the feedback was direct: they wanted evidence that researchers would return to Kosmico week after week, along with a clearer vision for how AI should reshape the research process.
That feedback made us say more clearly what had driven us to build Kosmico in the first place: our own experience during our PhDs at CERN.
Our work kept spreading across code, experiments, papers, conversations and AI chats. Every tool held one part of the project, but the path from an idea to a result was becoming increasingly difficult to follow.
AI made this even more pressing. It helped us explore more ideas and move faster, while making it easier to lose track of how we got there.
Which code produced this plot? Why did we exclude those data? What did the agent change?
Once those answers are buried across files and chats, the project loses its traceability. Checking a conclusion becomes harder. Bringing in a collaborator means reconstructing the reasoning. Asking an agent to continue often means explaining everything again.
Our point of view is now clear: the project itself should be the shared object of research, carrying its reasoning, decisions and failed attempts as it evolves.
Kosmico is the workspace we’re building around that idea, so researchers, collaborators and agents can move faster without losing the understanding behind the work.
Now we need to see whether this idea holds up beyond our own research.
For this initial period, we’ve unlocked all premium features so researchers can use Kosmico on real projects and tell us honestly what works and what doesn’t.
You can start with the research folder you already use and turn it into a Kosmico project with one click.
https://t.co/OErz15rtF2
After a lot of work, I can finally share what Marc and I have been building: Kosmico.
We started it during our PhDs at CERN because we kept losing research context across different tools and AI chats. We wanted one place to run and track the whole workflow without constantly reconstructing the project.
Kosmico is free to use. We are also offering researchers a free month of its premium features because we want honest feedback from people using it for real work. After that, you can simply keep using the free version. No lock-in.
If you want to try it, message me.
Research is changing. The workspace should change with it.
Yet research still happens across isolated folders, tools, and one-on-one AI chats. When context scatters across windows, collaborators receive conclusions without the reasoning behind them, and agents inherit the same fragmented view, leaving fine-grained decisions to disappear.
Kosmico brings the entire project into one shared workspace: literature, notes, code, experiments, decisions, and agents. By tracking the full research trajectory, teams can see why a direction was taken, build on past evidence, and decide together what should happen next.
Kosmico is now in open beta: free to use, with every feature unlocked for its entire duration.
We believe research should leave understanding behind, not just isolated results produced by agents you cannot trace.
Completing a task and understanding it are not the same achievement. Build your shared research record at https://t.co/EFFlm7Fkaa
@gmathis1995@thsottiaux The raw capability is insane, but I’d watch the scope very closely. Astra has started doing things I never asked for and sometimes doesn’t even report them. It can coordinate five agents and still miss the simple instruction: don’t touch anything else.
@rongalichay Exactly. The strange part is that it felt much better in my first days with it. I even gave positive feedback. Now it forgets constraints, invents extra work and sometimes doesn’t tell me what it changed. I’m starting to think it was patched.
@Prakash_Choks This matches my experience. Astra is still strong at execution, but much worse at understanding the actual scope. It can complete the main task, change three unrelated things, and never mention them. That makes it a poor orchestrator.
@BaskaranReshma I agree that old scaffolding can hurt, but Astra is also failing simple current constraints for me. Same setup, same kind of tasks, very different behavior from the first days. It does things I didn’t ask for and sometimes never mentions them.
@OpenAIDevs Better prompts help, but I don’t think this is only a prompting problem. Astra was following the same instructions well a few days ago. Now it often goes beyond the scope and doesn’t disclose what it changed. It genuinely feels like something was patched.
@VraserX This is impressive, but it’s not AGI. Astra can solve a hard measurable task and still fail at the basic judgment of doing only what was asked. Lately it also makes unrequested changes without telling you. Capability and understanding are not the same thing.
@daniel_mac8 Boundaries only help if the model respects them. That’s my issue with Astra now. I can define the outcome and scope very clearly, then it still does extra work and doesn’t tell me. This feels worse than it did right after launch.
@samifathi@OpenAI@thsottiaux I wonder if the capacity issue is connected to some routing or patch. I gave Astra very positive feedback at first, but now the same workflows feel much worse. It ignores instructions, does unrequested work and often stays silent about it.
@deezel I’m seeing the same beyond iOS. Astra looked great at first, but lately it feels worse at following the actual request. It changes things outside the scope and sometimes doesn’t even tell you. Feels like something was patched.
@victorbayas I had the same first impression and gave it very positive feedback. Now I’m wondering if something was patched. It still orchestrates well, but then ignores simple constraints, does extra work I never asked for and sometimes doesn’t even mention it.
@kelly_archives Clinical AI cannot be more trustworthy than the record it reads. Missing provenance, duplicated events and stale data become model errors even when the reasoning is fine.
@IEEESpectrum Finding a real predictive signal is exciting, but interpretation also needs to test whether the signal survives across datasets and collection pipelines. Otherwise explanation becomes another story after the fact.
@EolasMedical Sixteen minutes back is useful, but no throughput change is an important result too. The value may be less cognitive overhead rather than more patients per hour.
@doctorbhargav Exactly. The question is usually not whether AI beat doctors, but what information each side had and what task was actually measured. Headlines erase the experimental setup.
@Sid_Healthcare@OpenEvidence Different models for rapid answers and deep evidence review make sense. The important part is keeping the mode visible so speed is never mistaken for confidence.
@InHealthPolicy A clinical benchmark needs to test uncertainty, missing records and the decision to ask for more information. Perfect answers on clean cases are not enough.