@elliotarledge complete the cycle and make an AI that then extracts highlights from the gameplay and makes ig reels out of it, and another AI that watches the reels
GradCraft officially trained the first clanker
it is a prime slop generating machine
its a screen from README.md in my repo
finally proved gradcraft <3 finito
do u benchmark the kernels you write against production grade ones? Im writing flashattn rn, and i find it funny that the idea is pretty much the same, but if I compare my code to pytorch i am still significantly slower because of 100 diff intrinsics, tensor cores, math tricks/insights, etc. that pytorch uses to make it fast
I just wrapped up README.md for GradCraft
everyone on windows can now copy the code and train a clanker on their machine
is it better than pytorch? not even close. Is it my own? hell yea
link in bio
OH MY FUCKING GOD THE CLANKER I PRETRAINED IN GRADCRAFT CAN WRITE PYTHON!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
it is a bit retarded on the binary search because it inverted update of left adn right
BUT OH MY GOD IT IS VALID CODE
IT EVEN WORKS WITH DIFFERENT VARIABLE NAMES
I wrote an engine in C++ to train a GPT style LLM on my 3090, it was like 180M params, and today I decided to check inference on it, its like 40% training done
Its so stupid lmao, I learned to not just scrape github for small LLMs
PROMPT (pretrained so couldn't just ask it):
#include <string>
#include <algorithm>
// Optimal solution to reverse a string
std::string reverse_string(std::string s) {
Generating...
--- PROMPT---
#include <string>
#include <algorithm>
// Optimal solution to reverse a string
std::string reverse_string(std::string s) {
--- MODEL OUTPUT ---
nameInstr keyflagditionally { write_UnicodeString(keychar_cStrSQpack); false;}});
total_depth += step;
}
strides mu(min_batch, boundary_length* extent);
------
Later I wanted to print a buffer and got blasted with an unfinished game engine from 2018
@flatkey101 the amount of optimizations you can add to an autograd engine is absolutely diabolical, my v1 of the gpt training script was like 60x slower than it is now
calling cudaMalloc in a training loop vs using a memory pool gave a 2x difference. Article below <3
177ms/step vs 361ms/step
core is: with cudaMalloc the CPU first waits for the GPU to finish going math, and then the GPU waits for the CPU to finish allocating memory and queueing up the kernel.
https://t.co/sTpTWbjmdM
the 180M model came out absolutely retarded after 40% of training, so i started training on python instead of cosmopedia and C++. I think it will be easier for it to grasp and won't have to memorize 67 676 676 biological names from cosmopedia
Also it trains 4x faster. 90M params