A senior citizen,lost 20L+ in FD Fraud at Central Bank India. Bank staff was involved. FIR filed 10 months ago. Police minimized bank’s role. RBI closed the case citing 'sub judice' despite 0 court hearing.Accountability & action? @RBI@FinMinIndia@centralbank_in@PMOIndia
Announcing new aisuite capability: Easy function calling with LLMs! Function calling (tool use) is an important capability for agentic workflows and other LLM applications, but is cumbersome for developers to use (left column in image). Our open-source aisuite package simplifies it to just one command (right column), and works for multiple LLM providers.
Hope this makes implementing agents easier for developers, and thanks Rohit Prsad & team for working with me on this!
https://t.co/gwz9oKTCFx
We have to take the LLMs to school.
When you open any textbook, you'll see three major types of information:
1. Background information / exposition. The meat of the textbook that explains concepts. As you attend over it, your brain is training on that data. This is equivalent to pretraining, where the model is reading the internet and accumulating background knowledge.
2. Worked problems with solutions. These are concrete examples of how an expert solves problems. They are demonstrations to be imitated. This is equivalent to supervised finetuning, where the model is finetuning on "ideal responses" for an Assistant, written by humans.
3. Practice problems. These are prompts to the student, usually without the solution, but always with the final answer. There are usually many, many of these at the end of each chapter. They are prompting the student to learn by trial & error - they have to try a bunch of stuff to get to the right answer. This is equivalent to reinforcement learning.
We've subjected LLMs to a ton of 1 and 2, but 3 is a nascent, emerging frontier. When we're creating datasets for LLMs, it's no different from writing textbooks for them, with these 3 types of data. They have to read, and they have to practice.
DeepSeek (Chinese AI co) making it look easy today with an open weights release of a frontier-grade LLM trained on a joke of a budget (2048 GPUs for 2 months, $6M).
For reference, this level of capability is supposed to require clusters of closer to 16K GPUs, the ones being brought up today are more around 100K GPUs. E.g. Llama 3 405B used 30.8M GPU-hours, while DeepSeek-V3 looks to be a stronger model at only 2.8M GPU-hours (~11X less compute). If the model also passes vibe checks (e.g. LLM arena rankings are ongoing, my few quick tests went well so far) it will be a highly impressive display of research and engineering under resource constraints.
Does this mean you don't need large GPU clusters for frontier LLMs? No but you have to ensure that you're not wasteful with what you have, and this looks like a nice demonstration that there's still a lot to get through with both data and algorithms.
Very nice & detailed tech report too, reading through.
"Google Scanned
Objects (GSO) dataset,
a curated collection of over 1000
3D scanned common household items for use in the Ignition
Gazebo [7] and Bullet [8] simulators"
17 action figures, 28 bags, 254 shoes, and more!
https://t.co/ViOjkArLLF
I am over the moon to announce:
1) I'm now a professor at University of Queensland (UQ), the top institute in my home state!
2) I'll be teaching a brand new deep learning course at UQ from April, which will form the basis of a new @fastdotai course! 🧵
https://t.co/RAMaHb7eZ2
Informative thread on using cloud GPUs. My personal favorite is currently the T4. Offers a good bang-for-buck value, and it's still affordable to use instances with 4 of them for multi-GPU training.
For a work assignment I had to create a script which uses a neural network to detect sums on a piece of (rotated) paper and solve them. I'm pretty satisfied how it turned out
#Opencv#TensorFlow#machinelearning
I wanted to be able to quickly prototype an idea (such as a racetrack for a video game) by drawing it out on a piece of paper and using computer vision to translate it to something I can try out. Pretty satisfied with how it turned out.
#AugmentedReality#Computervision#gamedev
Augmented Reality : real-time recognition of handwritten math functions and drawing their graphs
Pretty satisfied with how my first prototype turned out. Next step, sin/cos/tan and integrals (and ignoring my hand).
#machinelearning#computervision#augmentedreality
Racing with SenseGlove! Our Computer Vision expert @FolmerMartijn always finds the most creative ways to improve #CV detection.
Curious to see how you can use SenseGlove for #VR training? https://t.co/sYDwCtfHx3
#computervision#motiontracking#training
DietGPU: fast specialized lossless compression on Nvidia GPUs
If you have slow network fabric, this can speedup distributed training by a lot.
Authored by Jeff Johnson (gh:wickedfoo) who wrote a lot of the PyTorch CUDA code and faiss-gpu.
https://t.co/eNvAJ4CJez
@heads0rtai1s Not sure if that's what you are looking for, but we had a tech-focused overview ("Machine Learning in Python: Main Developments and Technology Trends in Data Science, Machine Learning, and Artificial Intelligence") that summarizes a bit of the landscape https://t.co/U99TAutUzy
Opening & investing $100M into Ola Futurefoundry, our advanced engineering & design centre in U.K. Will work closely with team in Bangalore & focus on high perf. automotive engineering, vehicle design & battery tech. Can't wait to show you guys what we've been working on here!