@zimmskal Btw. literatur-empfehlung: https://t.co/Adj6Ss797G (das paper) bzw. siglip1.
Hab gerade für einen Kurs einen classifier für https://t.co/YM577iLTSr gebaut, und (nach data cleaning + 1x siglip) 86% auf img|text und 96% auf img+text mit einem ~10mb 3 layer MLP in <1 min train time.
@zimmskal Wirst du aber vmtl. semi easy in ein beliebiges Model rein-finetunen können, mit ein paar tausend svg's die du nach jpg convertest und dann in sowas wie gemma3 rein pastest und als expected out das svg nimmst. Idealerweise alle validen permutationen des svg-dom-trees.
@zimmskal Ist bottlenecked für alle llms die clip als image input encoder verwenden, weil das die location von elementen truncated.
Und weil das nicht wirklich in den trainingsdaten sein wird.
D.h. deepseek janus pro, und (afaik) das vor ~20h updated gpt-4o. Ggf. auch gemma3 mit siglip.
Lukewarm take: Catching up after not being on twitter for a month or so is impossible if you want to unfollow+refresh in-between, because sorting tweets by post-date would... make it feel less like a slot machine?
And ctrl+f is borken. Fuck this, kthxbye.
@Naehmo Ich find den Gen Z baggy style mit 40 immer noch bissi cringe, aber definitiv am sympathschisten/hottesten/nicht-robotischsten so far.
Dafür sind die talking points halt musk/trump light.
Who are your top 5 programmers of all time and why?
Mine:
1. Fabrice Bellard (ffmpeg, tinycc, quickjs)
2. John Carmack (doom, quake)
3. John McCarthy (lisp, father of AI, invented GC, timesharing)
4. Linus (linux man)
5. Dennis Ritchie (C and unix, K&R book)
The ultimate jailbreak for open LLMs: find the "direction" of the concept of refusal and substract it.
Boom, LLM can't refuse your requests and gives bomb tutorials... with soap, water, citrus and vinegar.
Nice data preprocessing, meta.
2/ Activation Ablation Augmentation
Research team that identified how to nullify activation layers in Llama3 responsible for censorship
@song_minjune @IanSears96@Nottlespike
🥉 3rd place
@zimmskal@HochreiterSepp It's just a model/architecture currently. The biggest it has been trained in the paper was to 1.5B capacity - with ~60 hellaswag score which is (at best) the same as an average 1.*B model. Aka bad.