@GergelyOrosz@can The notes on their site are amazing too. I remember studying the H/W and S/W side of CUDA from there. Very concise and great for starting.
This is helpful if you let your app run in background on Windows.
I recently found out that Windows deprioritizes your app (reclassifies the process to low Quality-of-Servie/QoS) when you minimize the app or move it to background. The OS re-schedules the threads from P-cores to E-cores.
This is something I couldn't see/prove using the good ol' Windows Task Manager - everything looks good there. I had to use xperf to profile my process and Windows Performance Analyser (WPA) to visualize/analyze the results. The video shows how this was achieved.
As a fix, we use `SetProcessInformation(GetCurrentProcess(), ProcessPowerThrottling, โฆ to opt-out of execution-speed throttling, pinning the process to High QoS regardless of window state.
The second image shows the improvement - QoS stays as "Important" even when the app is minimized.
#windows #debugging #QoS
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
@hardmaru Exactly! I have been using 3.1 Pro for most of my tasks. I have access to Fable but never used it. Maybe its because I dont' let my agent read a doc and build/test entire thing from scratch. And rather work with it iteratively. So older models like 3.1 Pro and Opus4.6 suffice.
This Independence Day ๐ฎ๐ณ, I asked, "What will the next century of young Indians look like?"
$0.01 Drone Deliveries. UAVs that never land. Groceries in under 10 minutes.
Meet The 22nd Century Indian. A Documentary on a New India.
I am looking to hire a builder (max 2, if I am very impressed, we don't need more than that, trust me) in my team to work alongside with me.
Builders = Problem solvers
Problems from every domain, be you are comfortable or not.
What I can promise:
> Complete autonomy on research and innovation
> You will be challenged everyday (from me)
> And goes without saying "unlimited tokens/ unlimited tools"
Other formalities can be cleared by HR but its lucrative
What I need from you all to pitch yourself:
> Brainstorm/research an actual problem that Razorpay/RazorpayX is facing (talk to working folks in rzp/rzpX to get more context) and prepare a technical solution for that (don't dm me asking for problems, duh)
> Show something crazy you have built (not those ones hosted on vercel please).
startup founders are encouraged to pitch their company.
> Wildcard (anything that you think can win me over)
Brownie points : I like going deep into the technicalities and seeing how your product holds in extreme case.
Be ambitious and fearless!
Apply here : https://t.co/k2MPXpceAR
code auto-regressive decoding today!
I just wanted to see KV Cache in action - the simplest version possible. But it had to be part of a decoder and as bonus, I got to see how causal masking looks with and without KV-cache.
Also, the full forward pass through a decoder seemed so simple first but see the pic and you'll agree that coding even the simplest transformer is so different from conventional CNN code we knew.
But had fun; highly recommend everyone to do this to have a memory map.
#autoregressive #decoding for #kvcache #llm
I don't know why nsight-systems (nsys) is so popular while I never heard of nsight-compute (ncu) until last week. Only when I wanted to *see* how the shared memory reduced global-memory reads.
Look at the pics below
#nvidia ncu > nsys
For CUDA people who've had hard time visualizing the threads working before and after using share memory, I have tried to share an example here: https://t.co/sLNuLPjBlQ
#shared memory + #multithreading for #CUDA
One of the best pieces on preparing for ML roles.
Note that this might not be best for agent-based roles or RAG/MCP but apart from that, very much recommended.
Adding a snippet from the same >)
https://t.co/vjRxgf7oQf
@karmakarthik Isn't this too much? AI is very segregated nowadays. It has to be Computer Vision, NLP, purely agentification of existing stack, etc. What you mentioned is mostly relevant to agentification part where backend is needed too. Be specific or burn out.
I feel that this is so much relevant to the tech development happening right now. Every student/engineer wants to be a millionaire in the next 3 years.
Looking for people who are enjoying the journey, rather than rushing it.
I think @karpathy set a good example.