@dhillon_p Actually I tried using OCI and the security guard in Bangalore let me through but in Bhopal he said he needs the passport - so I have been using the passport since :)
I think OCI is not a general purpose ID, also see here: https://t.co/yfxERQbk45
While taking a domestic flight in India, I had a genuine question of whether I can use my OCI (overseas citizen of India) card as a valid proof of ID, and so I googled it and got the following conflicting results.
Factuality remains a core problem to fix.
🚀 Excited to share the research I worked on during my summer internship at @GoogleAI! We developed FRAMES (Factuality, Retrieval, And reasoning MEasurement Set), a challenging high-quality benchmark for evaluating retrieval-augmented large language models. FRAMES tests LLMs on retrieving relevant info, reasoning across documents, and providing factual responses to complex questions. 🧵👇 #AI #RAG #Eval
Dataset link: https://t.co/krEijwOist
Paper link: https://t.co/WVyV2izxZk
Ever wondered if model merging works at scale? Maybe the benefits wear off for bigger models?
Maybe you considered using model merging for post-training of your large model but not sure if it generalizes well?
cc: @GoogleAI@GoogleDeepMind @uncnlp
🧵👇
Excited to announce my internship work on large-scale model merging! We explore what happens when you combine larger and larger language models (up to 64B parameters!) and how different factors –model size, base model quality, merging methods, and # of experts– impact held-in performance and generalization.
📰: https://t.co/hIVBFh9Rb5
Kaling: The real reason I'm here is that deep down, I truly believe that as a woman of color and a single mother of three, it is incredibly important that I be appointed ambassador to Italy.
It's delightful that we are #1 here, but...
We can be #3 in the next few weeks, and then again become #1 sometime
What's really exciting about working on Gemini is that we will bring the best user-experience in our products through these models, these metrics are by-products.😉
Exciting News from Chatbot Arena!
@GoogleDeepMind's new Gemini 1.5 Pro (Experimental 0801) has been tested in Arena for the past week, gathering over 12K community votes.
For the first time, Google Gemini has claimed the #1 spot, surpassing GPT-4o/Claude-3.5 with an impressive score of 1300 (!), and also achieving #1 on our Vision Leaderboard.
Gemini 1.5 Pro (0801) excels in multi-lingual tasks and delivers robust performance in technical areas like Math, Hard Prompts, and Coding.
Huge congrats to @GoogleDeepMind on this remarkable milestone!
Gemini (0801) Category Rankings:
- Overall: #1
- Math: #1-3
- Instruction-Following: #1-2
- Coding: #3-5
- Hard Prompts (English): #2-5
Come try the model and let us know your feedback!
More analysis below👇
On behalf of the American people, I thank Joe Biden for his extraordinary leadership as President of the United States and for his decades of service to our country.
I am honored to have the President’s endorsement and my intention is to earn and win this nomination.
Hard for me to support someone with no values, lies, cheats, rapes, demeans women, hates immigrants like me. He may cut my taxes or reduce some regulation but that is no reason to accept depravity in his personal values. Do you want President who will set back climate by a decade in his first year? Do you want his example for your kids as values?
Google announces Leave No Context Behind
Efficient Infinite Context Transformers with Infini-attention
This work introduces an efficient method to scale Transformer-based Large Language Models (LLMs) to infinitely long inputs with bounded memory and computation. A key
Google presents Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
1B model that was fine-tuned on up to 5K sequence length passkey instances solves the 1M length problem
https://t.co/zyHMt3inhi
Bard with Gemini Pro is now available in over 40 languages and 230+ countries and territories, bringing its top 2 most preferred status across the world (👀 @lmsysorg).
AND you can now bring your imagination to life with image generation. It's optimized for speed and is available in English across most countries... let your imagination run wild, at no cost.
AND we’re also making it easier to corroborate Bard’s responses. We’re expanding our Double Check feature, which is already used by millions of people in English, to more than 40 languages across the globe.
Thank you, PaLM2 for all you did to get us here 🫡. Rest easy.
Ok back to work... read more here: https://t.co/UMCrjbdGsf
🔥Breaking News from Arena
Google's Bard has just made a stunning leap, surpassing GPT-4 to the SECOND SPOT on the leaderboard! Big congrats to @Google for the remarkable achievement!
The race is heating up like never before! Super excited to see what's next for Bard + Gemini Ultra release.
(this will be the last response just for the record; this type of engagement is not why I use this app)
1. Dataset was released along with the paper. again, eval on a dataset of this scale really doesn't take long, especially for google
2. this was an one-off project crowdsourced by a group of researchers of similar interests, not a well-planned-ahead-of-time project funded by any agency or big tech.
Hi @emilymbender, I'm one of the lead authors of MMMU. I can certify that 1) Google didn't fund this work, and 2) Google didn't have early access. They really like the benchmark after our release and worked very hard to get the results. It doesn't take that long to eval on a dataset