https://t.co/8ZHBAkFlT8
🚀 321 New Sora 2 Invite Codes Available! 🚀 SAHN HOVA
To claim yours:
1️⃣ Like ❤️ this post
2️⃣ Repost 🔁 to share the news
3️⃣ Comment “SAHN HOVA ” 💬
(Make sure you’re following, or I can’t DM you your code.)
#Sora2#AI#InviteCode#Giveaway#AIVideo #OpenAI #SahnHova #birdtalk
My first @GoogleDeepMind project: How do LLMs recall facts?
Early MLP layers act as a lookup table, with significant superposition! They recognise entities and produce their attributes as directions. We suggest viewing fact recall as a black box making "multi-token embeddings”
GPT-4V is a Generalist Web Agent, if Grounded
Similar to LLMs, there is going to be a huge exploration space with large multimodal models (LMMs).
The big question is what can LMMs like GPT-4V and Gemini do beyond traditional tasks like image captioning and visual QA?
This new exciting work from Zheng et al. explores the potential of GPT-4V as a generalist web agent. In particular, can such a model follow natural language instructions to complete tasks on a website?
The authors first developed a tool to enable web agents to run on live websites.
Findings suggest that GPT-4V can complete 50% of tasks on live websites. This is possible through manual grounding of its textual plans into actions on the websites.
The key challenge here is the grounding part. For grounding, this work leverages both HTML text and visual elements which is unique for websites compared to natural images.
GPT-4V exhibits long-range planning, webpage content reasoning, and error correction capabilities among other things.
There is still a lot work of needed to improve grounding and reduce hallucinations for such LMM-powered web agents. But this is a great start demonstrating the huge potential of LMMs as web agents.
paper: https://t.co/7GdWIqn591
code: https://t.co/2X8YgTGto5