#WWDC25 is tomorrow and here is my wishlist:
- Small Language Models for devs to use on device
- A library for the SLMs to be agentic (Model, Instructions, Tools)
- A routing library for those SLMs to orchestrate them
- A cloud option to run tasks async
- Agents in Xcode
for our next open source project, would it be more useful to do an o3-mini level model that is pretty small but still needs to run on GPUs, or the best phone-sized model we can do?
WOW! Apple just dropped Core ML optimised models for FastVIT, DepthAnything & DETR 🔥
> Quantised models for Image Classification, Monocular Depth Estimation, Semantic Segmentation
> Along with the checkpoints, they report detailed benchmarks on Inference speed, accuracy and more
> Bonus: We ship sample apps to get started with these models quickly!
The time for on-device AI is now! 🤗
With #WWDC24 less than a week away, here is my wishlist:
- On-device inference with a Small Language Model
- SwiftData with vector embeddings
- LoRAs for on-device models
- A framework for multi agent systems
- Local copilot for Xcode
We’re releasing GPT-4 — a large multimodal model (image & text in, text out) which is a significant advance in both capability and alignment.
Still limited in many ways, but passes many qualification benchmarks like the bar exam & AP Calculus: https://t.co/L6VGJ0WfFv
2/2 I also think this role would require encyclopedic domain knowledge. Like a living index of stuff related to a specific field or task. So the person can ask the right prompt with proper descriptions and know enough to check the answer to re-engineer the question to fix it.
1/2 With StableDiffusion and ChatGPT being so reliant on good prompts and queries, I’m starting to think that there might be a new career for people that ask good, inquisitive questions. What job title do you think it’ll be? #stablediffusion#ChatGPT#dalle2#midjourneyAi