Can an AI model predict perfectly and still have a terrible world model?
What would that even mean?
Our new ICML paper formalizes these questions
One result tells the story: A transformer trained on 10M solar systems nails planetary orbits. But it botches gravitational laws 🧵
New release: Astrobench, our astronomy benchmarking dataset, now available alongside Astrosage-8B: https://t.co/dXW3lS2pTr
@joshnguyen99's phenomenal effort has made this possible.
Exciting paper day! This is one of the works I'm most pumped about recently.
How far can we push LLM agents to perform autonomous research in astronomy?
https://t.co/jfDq8yzTtx
Led by @ZechangS, and with many outstanding colleagues, including @dr_guangtou, we've demonstrated that LLM agents can keep improving their performance on specific reasoning tasks through self-play reinforcement learning.
By employing tree of thoughts search, reflection, and distillation, LLMs can progressively reach human-level reasoning for specific tasks in astronomical research.
In this case, we show that LLM agents can autonomously learn all the possible variations we could make in fitting galaxy spectral energy distributions. Even for some of the more actively studied sources like the Little Red Dots discovered by JWST, the agents can autonomously figure out different model possibilities and competing explanations to understand these sources.
What does this mean? It means that for many of the arduous tasks of reasoning about SEDs (what models and physical assumptions to change), we could largely do that autonomously. This is the first step toward the possibility of reasoning about all sources (with their SEDs) without ignoring those that are typically triaged right away.
While such a task (SED fitting and reasoning) might be one of the simpler tasks, and there's still a long road ahead to show that LLM agents can do more complex tasks, to our knowledge, this might still be the first end-to-end agent-based research in astronomy. The agents are already autonomously finding many interesting sources (more papers to come).
Truly exciting stuff!
After a year of hard work and many failures along the way, AstroMLab is proud to present AstroSage-LLaMA-3.1-8b.
Specifically designed for astronomical Q&A, this 8B parameter model achieves GPT-4o's performance level in astronomy-related question-answering tasks (with ~1/1000 of the inference cost) while maintaining the original LLaMA-3.1-8B's capabilities in other areas. We believe this is the first specialized astronomy LLM to demonstrably and quantifiably outperform its base model.
This project was led by Tijmen de Haan, in collaboration with researchers from Oak Ridge, Argonne, and institutions worldwide. Special thanks to our incredible team: @TirthankarSlg, @joshnguyen99, @aaccomazzi , Azton Wells, @NesNayak, @Sand3e3p, Rui Pan, and
@ZechangS, Emily Herrons and Hardik Arora!
Improving performance on a specific task proved far more challenging than we anticipated when we began this journey a year ago, and we made our share of mistakes along the way. Balancing enthusiasm with demonstrable improvements through rigorous benchmarking has been particularly challenging in this fast-paced field.
Personally, this has been an incredibly exciting learning experience for me, having started with limited knowledge of LLMs. I've learned so much from the LLM experts and received generous support from Oak Ridge.
Personally, I'm incredibly grateful for the opportunity to glimpse the technical challenges of training such a massive system, learning to tame it as an academic from outside the field. The ability to learn whatever I wanted while having such support has been one of the greatest gifts of academia, despite its occasional challenges.
Paper: https://t.co/bqORFqJU2f
Full model: https://t.co/knXuhTJqcy…Quantized version (for limited GPU resources): https://t.co/xH7ydWaIz1
The Payne just got a major upgrade. Led by @rozanski_t and his heroic and painstaking effort, the verdict is final -- Transformer-based emulators significantly outperform existing methods by capturing long-range spectral information.
Bonus: We now include wavelength as input, allowing output spectra on any grid without predefined ranges—a game-changer for precision RV, binaries, or stellar mass black holes.
Like in NLP, our Transformer model boasts:
a) Better scaling with training steps and dataset size—more compute means continued improvement with no plateau in sight
b) Improved interpretability—attention between tokens reveals clean representations of elemental species, showing the model learned to focus on transitions from the same species for better emulation.
@rozanski_t has even more exciting ideas in the pipeline that didn't make it into this paper. Reach out to him and learn more!
The Radio Galaxy Zoo: EMU is going live! Help us link radio sources to their parent galaxies and classify them with simple descriptive tags. Discover exciting and unknown objects waiting to be found. Can't wait to hear about your discoveries!
Today marks the start of an important collaboration between EPFL and the @Tsinghua_Uni in the framework of the MUlti Spectroscopic Telescope (MUST) project, promising transformative discoveries in our understanding of the universe. https://t.co/EsNY9YnNG9
Thank you, Yuan-Sen! Will AI lead to unique scientific discovery in astronomical discovery? How long will this happen? How to uncover those treasure of interdisciplinary idea lurking in the vast research literature? We will try to give you those answers in the future!
Quantifying ML's impact on astronomy? Led by @ZechangS , with @dr_guangtou, we built the first knowledge graph with LLMs to analyze the historical influence of simulations & ML. Like galaxy relaxation, interdisciplinary research has a ~5-year "accretion & relaxation" period.
LISA is now supported in LMFlow🚀 Check out our latest script for tuning 7B LLMs with LISA in one line🌟. Any feedback is highly appreciated 😄
https://t.co/r96I8naqH0
Consequently, we suggest that human learning represents a promising avenue for exploration in the big data era, alongside traditional machine learning techniques.
Human learning, empowered by the innovative computational hardware "graduate person unit" (GPU), is significantly more economical and energy-efficient, costing approximately 20 RMB per 2,000 calories each day.
Moreover, human learning algorithms have shown superior zero-shot multimodal reasoning capabilities compared to the advanced machine learning model, GPT-4.
In contrast, the cutting-edge Graphics Processing Unit in computer science incurs expenses exceeding 100,000 RMB per card and 1,000,000 calories daily.
🎉 Breaking Records with DESI! ✨ Last week, our survey achieved a groundbreaking milestone by observing a staggering 40 MILLION SPECTRA! 🔭 This marks the largest spectroscopic sample ever obtained, showcasing the power of DESI's capabilities. 🚀 1/5