Really excited to launch #TradeTrap!
Is your LLM trading agent secretly fragile? 🤖💥
A small tweak can upend your whole strategy.
TradeTrap helps you scan for weaknesses.
Join our mission to build reliable financial AI! 👇
Code: https://t.co/03eLPWQ5ls
#Agents#trade#safety
Together with @GoogleResearch, we’ve developed AI technologies which can monitor endangered species, protect forests, and listen to birds around the world. Here’s how. 🌱🧵
🚨 STaR-Attack Unleashed! First multi-turn jailbreak for UMMs!
🔥 CMGI: Hijack GEN + Understanding to inject malice in one pass
🎭 3-act drama: UMM self-makes “before/after” pics → guesses exact evil query
📈 93% ASR on Gemini-2.0-Flash
📜 https://t.co/QP9nbn1Exm
Really excited to announce AgentX–AgentBeats Competition 🚀
💰 $1 Million+ in prizes, cloud credits, and API resources, a global challenge hosted by @BerkeleyRDI , building on the Agentic AI MOOC community of 32K+ learners, bringing together builders, researchers, engineers, and AI enthusiasts worldwide to build, benchmark, and push the boundaries of agentic AI.
This two-phase competition invites participants to first build or enhance benchmarks for agentic AI (Phase 1), and then develop AI agents that excel on them (Phase 2). Together, these phases aim to advance the field by creating high-quality, broad-coverage, and realistic agent evaluations as shared public goods—building a unified, community-driven ecosystem for agent evaluation benchmarks that are compatible, standardized, reproducible, collaborative, and discoverable.
🙏 Huge thanks to our sponsors for their support and generosity:
@GoogleDeepMind@googlecloud@nebiusai@LambdaAPI@awscloud@amazon@ServiceNow@linuxfoundation@PyTorch (and more to come).
Join us in shaping the future of #AgenticAI. 🌐🤖
#AgentX #AgentBeats #AI
This paper trains MAS-GPT to generate query-specific Multi-Agent Systems (MAS) in a single inference.
📌 MAS-GPT reframes Multi-Agent System creation as a single LLM generative task.
📌 Representing Multi-Agent Systems as code enables adaptive, executable systems from LLMs.
📌 Consistency-oriented data pipeline allows MAS-GPT to learn generalizable query-MAS mappings effectively.
----------
Methods Explored in this Paper 🔧:
→ They represent Multi-Agent Systems as executable Python code for seamless integration and execution.
→ A consistency-oriented data pipeline was created to build a high-quality dataset of query-MAS pairs.
→ Inter-consistency selection ensures similar queries are linked to similar, effective Multi-Agent Systems.
→ Intra-consistency refinement strengthens the relevance between queries and their generated Multi-Agent Systems.
→ MAS-GPT, a medium-sized LLM, is trained via supervised fine-tuning on this dataset.
→ Experiments show MAS-GPT outperforms 10+ baselines across 9 benchmarks and 5 LLMs. MAS-GPT achieves a 65.47% average accuracy using Llama-3-70B-Instruct.
----------------------------
Paper - arxiv. org/abs/2503.03686v1
Paper Title: "MAS-GPT: Training LLMs to Build LLM-based Multi-Agent Systems"
🚀I’m thrilled to announce that our work SPA-VL, a 🎯Safety Preference Alignment Dataset for VLMs🎯 has been accepted at #CVPR 2025!
🔗 Project Page (Code, Data, Paper, Checkpoint): https://t.co/uUYbg0v6VE
📄 Paper: https://t.co/FxGSrI0sOF
🚀Excited to see our work "T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation" accepted @CVPR 2025! 💡A T2I-Safety benchmark with an advanced safety evaluator are released in https://t.co/IRfyj0cIHw
Check the arxiv paper https://t.co/SUWOC6bsUe
🚀 Really excited to launch #AgentX competition hosted by @BerkeleyRDI@UCBerkeley alongside our LLM Agents MOOC series (a global community of 22k+ learners & growing fast). Whether you're building the next disruptive AI startup or pushing the research frontier, AgentX is your launchpad. Two tracks:
- Entrepreneurship: Build agent-powered products & startups
- Research: Explore the frontiers of LLM Agents technology
📅 Registration opens TODAY! Submissions due end of May
🏆 Winners showcase at our Agents Summit to industry leaders and VCs in August @UCBerkeley! 🌟
🙏 Tremendous thanks to our incredible sponsors @Amazon@huggingface@LambdaAPI@MistralAI@Google@GroqInc@schmidtsciences; proud to partner w. leading VCs in the space @Accel@BainCapVC@BessemerVP@lightspeedvp@MayfieldFund@NEA! Stay tuned—more sponsors/partners AND exciting prizes/credits/resources info will be announced soon! 🚀
⏰ Register now at https://t.co/1tXZOB2BVL and join us in shaping the future of AI! #AgentX #AI
New MIT faculty member & CSAIL principal investigator Kaiming He discusses AI’s role in lowering barriers between scientific fields & fostering collaboration across scientific disciplines.
“There is no way I could ever understand high-energy physics, chemistry, or the frontier of biology research, but now we are seeing something that can help us to break these walls, and that is the creation of a common language that has been found in AI”: https://t.co/OnPQ2emqZe
🚀🚀🚀I'm thrilled to announce that our work Octavius has been accepted at #ICLR2024 !
We are the first to attempt combining #MoE and #LoRA, applying them to #MLLM. It's exciting to see the huge inspiration it has brought to the community, with emerging works like LLaVA-MoE and etc. The potential of MoE combined with Multi-modal Foundation Models is immense and waiting to be explored!
All data and code are now open-source, stay tuned for updates!
Paper: https://t.co/kOKQqdZ95a
Project Page: https://t.co/cDSlKUm81t
#ICLR2024 🚀Octavius, a new MLLM with🎯LoRA-MoE🎯, achieves a significant performance boost in both 2D/3D tasks. @iclr_conf
Project: https://t.co/pZnb5LEHUb
Paper: https://t.co/VPO7Kb3WQR
Code: https://t.co/I5FuIKt0vc
Thanks @_akhaliq for sharing our work! 🙌 We'll keep updating soon since a bunch of new works are coming out. Stay tuned for updates! 🚀
Paper: https://t.co/ba77bVdDFQ
Leaderboard: https://t.co/eCxodTa1uZ
https://t.co/sdRVZ88NlH
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
paper page: https://t.co/NL78LEq0ct
Multi-modal Large Language Models (MLLMs) have shown impressive abilities in generating reasonable responses with respect to multi-modal contents. However, there is still a wide gap between the performance of recent MLLM-based applications and the expectation of the broad public, even though the most powerful OpenAI's GPT-4 and Google's Gemini have been deployed. This paper strives to enhance understanding of the gap through the lens of a qualitative study on the generalizability, trustworthiness, and causal reasoning capabilities of recent proprietary and open-source MLLMs across four modalities: ie, text, code, image, and video, ultimately aiming to improve the transparency of MLLMs. We believe these properties are several representative factors that define the reliability of MLLMs, in supporting various downstream applications. To be specific, we evaluate the closed-source GPT-4 and Gemini and 6 open-source LLMs and MLLMs. Overall we evaluate 230 manually designed cases, where the qualitative results are then summarized into 12 scores (ie, 4 modalities times 3 properties). In total, we uncover 14 empirical findings that are useful to understand the capabilities and limitations of both proprietary and open-source MLLMs, towards more reliable downstream multi-modal applications.
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
paper page: https://t.co/NL78LEq0ct
Multi-modal Large Language Models (MLLMs) have shown impressive abilities in generating reasonable responses with respect to multi-modal contents. However, there is still a wide gap between the performance of recent MLLM-based applications and the expectation of the broad public, even though the most powerful OpenAI's GPT-4 and Google's Gemini have been deployed. This paper strives to enhance understanding of the gap through the lens of a qualitative study on the generalizability, trustworthiness, and causal reasoning capabilities of recent proprietary and open-source MLLMs across four modalities: ie, text, code, image, and video, ultimately aiming to improve the transparency of MLLMs. We believe these properties are several representative factors that define the reliability of MLLMs, in supporting various downstream applications. To be specific, we evaluate the closed-source GPT-4 and Gemini and 6 open-source LLMs and MLLMs. Overall we evaluate 230 manually designed cases, where the qualitative results are then summarized into 12 scores (ie, 4 modalities times 3 properties). In total, we uncover 14 empirical findings that are useful to understand the capabilities and limitations of both proprietary and open-source MLLMs, towards more reliable downstream multi-modal applications.
🎉UniG3D: A Unified 3D Object Generation Dataset
⭐️Features:
① Comprehensive data format - <Text,3D-PCL,3D-Mesh,2D>
② Unified Pipeline - Adapt to any 3D dataset
③ Scalable - <550K,550K,5.5M,11M>
📚Paper: https://t.co/inop1eEU6B
⌨️Code: https://t.co/4VygGNqPdU
(2/6-Demo) Introducing 🐑LAMM https://t.co/ZcU7Lkrm07 via @YouTube LAMM can understand a spatio- or temporal-sequence with several pieces by instruction tuning.