I'm playing with AI Agents today, using AutoGen.
After some initial hurdles, using Git as a source of communication seems the most promising path at present. The passing of context is reduced.. historical decisions are naturally persisted.. and it fits well with writing software.
There are two ways to start the day; shock or solace.
Shock - cold shower, hard exercise, ginger hot shots.
Solace - warm shower, gentle movement, coffee or tea.
One kickstarts the day. One warms it up gently.
Which one is best?
Which story do you want to live?
New Anthropic research paper: Scaling Monosemanticity.
The first ever detailed look inside a leading large language model.
Read the blog post here: https://t.co/6RYwxt6nWI
Microsoft's #Phi#AI models are interesting in many ways.
1 - Their training dataset contains a significant amount of synthetic data.
Copyright concerns plague other models, but if training data is generated - does that eliminate a whole host of concerns?
I've only just realised that M$'s #phi3 model has been named as a Small Language Model (SLM)
It feels like a branding-error.
Perhaps a better name: #ConciseLanguageModel ?
Do we have a name for models only trained on highly-curated and high-quality data?
@natfriedman Quiet, gardening robots would be the best.
I'd like one that quietly roams the local area picking up pieces of trash blown away on the wind.
I was thinking of getting a nice rug, a large TV, and a sofa in my study.
Now I'm thinking I should keep the space ready for a Holotile floor, and an Apple Vision headset.
I’m very excited to share our work on Gemini today! Gemini is a family of multimodal models that demonstrate really strong capabilities across the image, audio, video, and text domains. Our most-capable model, Gemini Ultra, advances the state of the art in 30 of 32 benchmarks, including 10 of 12 popular text and reasoning benchmarks, 9 of 9 image understanding benchmarks, 6 of 6 video understanding benchmarks, and 5 of 5 speech recognition and speech translation benchmarks. Gemini Ultra is the first model to achieve human-expert performance on MMLU across 57 subjects with a score above 90%. It also achieves a new state-of-the-art score of 62.4% on the new MMMU multimodal reasoning benchmark, outperforming the previous best model by more than 5 percentage points.
Gemini was built by an awesome team of people from @GoogleDeepMind, @GoogleResearch, and elsewhere at @Google, and is one of the largest science and engineering efforts we’ve ever undertaken. As one of the two overall technical leads of the Gemini effort, along with my colleague @OriolVinyalsML, I am incredibly proud of the whole team, and we’re so excited to be sharing our work with you today!
There’s quite a lot of different material about Gemini available, starting with:
Main blog post: https://t.co/NzSycJl7aE
60-page technical report authored by th Gemini Team: https://t.co/CEdMRyYSLo
In this thread, I’ll walk you through some of the highlights.