I just used 3,000 GPU-hours to test all 9 new OpenAI Whisper speech recognition models, two Talon acoustic models, and NVIDIA Nemo large and xlarge models. Whisper has a peculiar failure case. Here's what I think.
I wrote a really simple RISC-V (rv32i) JIT for x86_64 designed for gathering some stats for my upcoming Bluehat IL talk. It runs at about ~1 RISC-V instruction per 2 x86 cycles, and can create and run hello world ~8.6 million times per second on 96 cores! https://t.co/6UAo283Ngq
Cache Associativity can be surprising ¯\_(ツ)_/¯
Example in C# .NET
If you want a good explanation, then read @sergey_slotin post about this and many other problems here: https://t.co/Ub4k38Depx
#dotnet
It's a little surreal to see that all of these companies (and more) now have licenses for the Slug Library. A single idea that I had five years ago, and a lot of hard work on top of it, has changed the course of my career.
https://t.co/5jJEnAqYvg
Puzzled why a yara rule did or didn't match?
Let me introduce https://t.co/cY3G5MeOk6, a web-based #yara#debugger!
With #YaraDbg, you can see the:
1⃣ evaluation steps
2⃣ matched strings
3⃣ relationship among the rules
Okay @gamozolabs just blew my mind with this knowledge that x86 is an octal machine. How is this not more commonly understood. The opcode mods use values that are obvious enums when you see them displayed as octal.
https://t.co/NFjSzNAJsW
Using Vulkan Memory Allocator? I just pushed a big feature: a new API for selecting optimal memory type. I now recommend to just use allocCreateInfo.usage = VMA_MEMORY_USAGE_AUTO;
Legacy x86 fun fact: A far call in 64-bit mode can take a 32-bit and 64-bit offset in the far pointer on Intel (selected by REX prefix). AMD ignores the REX prefix and only supports 32-bit offsets. I'm currently checking the AMD manual to see whether this is documented...
So I did a thing.
Back a couple years ago, people were rewriting the classic 'wc' program (word-count) in their favorite programming language to prove theirs could be as fast as C.
So I decided to rewrite using my favorite algorithm instead: a "state machine parser".
New research: What if I told you that you could execute arbitrary code at runtime without allocating or writing to executable memory? Read more 👇
https://t.co/BOGVOOZUM6
I've completely rewritten a blog post detailing the differences between shader languages. What are some of the differences between WebGPU's shading language and HLSL, GLSL, and MSL? How do compute shaders compare across languages?
https://t.co/gFbgDAMDP2
I wish I had known about this a couple years ago: A really detailed, interactive diagram of WebGL’s internal, global state object, where you can see how each WebGL API call affects said object.
https://t.co/iKCzbvRUil
There are differences between AMD and Intel in their 64-bit page table format. Specifically the G(lobal) bit is ignored on Intel in the PML4, but not on AMD where it has to be zero. That kept me from switching to Long Mode on my AMD box for the kernel entry benchmarks... #x86
With inflation at 7.5%, you lose half your money in 9 years. The only way to outperform that consistently, that I have found, is crypto. Just this year I’ve already lost half my money.