@nCompass_tech is a performance optimization dev tool that unifies ๐ฝ๐ฟ๐ผ๐ณ๐ถ๐น๐ถ๐ป๐ด, ๐๐ฟ๐ฎ๐ฐ๐ฒ ๐ฐ๐ผ๐น๐น๐ฎ๐ฏ๐ผ๐ฟ๐ฎ๐๐ถ๐ผ๐ป and ๐๐ฟ๐ฎ๐ฐ๐ฒ ๐ฎ๐ป๐ฎ๐น๐๐๐ถ๐ of AI systems.
We automate the process of identifying and fixing the root cause of performance bottlenecks across all levels of the AI infrastructure stack. Our agent automatically surfaces key bottlenecks and works in tandem with Claude Code / Cursor to implement the fixes.
Using our tool, we were able to automatically improve a Hopper GEMM kernel taking it from ๐ฏ๐ฌ% ๐๐น๐ผ๐๐ฒ๐ฟ ๐๐ผ ๐ฏ% ๐ณ๐ฎ๐๐๐ฒ๐ฟ than a near optimal CUTLASS kernel within a day - this took us months before!
Weโre excited to officially release v0.1.0 of our VSCode extension that includes ๐ป๐ฎ๐๐ถ๐๐ฒ ๐๐ฟ๐ฎ๐ฐ๐ฒ ๐๐ถ๐ฒ๐๐ถ๐ป๐ด, ๐๐ฟ๐ฎ๐ฐ๐ฒ ๐ฑ๐ถ๐ณ๐ณ๐, ๐๐ฟ๐ฎ๐ฐ๐ฒ ๐ฐ๐ผ๐น๐น๐ฎ๐ฏ๐ผ๐ฟ๐ฎ๐๐ถ๐ผ๐ป and integration of our ๐๐ ๐ฎ๐ด๐ฒ๐ป๐ directly into Cursor / Claude Code. Weโre working closely with an initial set of users including the core maintainers of @vllm_project to build features that best support the community.
๐๐ณ ๐๐ผ๐โ๐ฟ๐ฒ ๐ฎ๐ป ๐ฒ๐ ๐ฝ๐ฒ๐ฟ๐ ๐๐ฃ๐จ ๐ฝ๐ฒ๐ฟ๐ณ๐ผ๐ฟ๐บ๐ฎ๐ป๐ฐ๐ฒ ๐ผ๐ฝ๐๐ถ๐บ๐ถ๐๐ฎ๐๐ถ๐ผ๐ป ๐ฒ๐ป๐ด๐ถ๐ป๐ฒ๐ฒ๐ฟ - bring your time spent optimizing performance down to days rather than weeks - view, collaborate and analyze faster than you do today.
๐๐ณ ๐๐ผ๐โ๐ฟ๐ฒ ๐ป๐ฒ๐ ๐ฎ๐ป๐ฑ ๐๐ฎ๐ป๐ ๐๐ผ ๐น๐ฒ๐ฎ๐ฟ๐ป ๐ต๐ผ๐ ๐๐ผ ๐ผ๐ฝ๐๐ถ๐บ๐ถ๐๐ฒ ๐ฝ๐ฒ๐ฟ๐ณ๐ผ๐ฟ๐บ๐ฎ๐ป๐ฐ๐ฒ - you now have an expert pair-programmer to consult at any time.
๐๐ฒ๐ ๐๐๐ฎ๐ฟ๐๐ฒ๐ฑ ๐ณ๐ผ๐ฟ ๐ณ๐ฟ๐ฒ๐ฒ ๐๐ผ๐ฑ๐ฎ๐ using VSCode, Cursor or Claude Code - https://t.co/lWXuYeoavi
We also work directly with you to identify opportunities for performance improvements. If you would like us to ๐ฎ๐๐ฑ๐ถ๐ ๐๐ผ๐๐ฟ ๐๐๐ฎ๐ฐ๐ธ and suggest ways to improve performance - reach out at [email protected]!
Thanks! Yeah that's a great point. I agree with your point when you're optimizing arbitrary codebases, but I think given our current focus around AI model inference/training execution and the GPU kernels that get launched, the space for optimization is a bit more constrained - which makes our lives a little bit easier.
But this can still arise for instance when you say you want to optimize a kernel but you don't specify the input sizes you want to target. Something like a split-k optimization is only really valid if you're running a small input, so you wouldn't want to suggest that for large batch size kernels. We're building our agent to be aware of these trade-offs so that it asks for the constraints you are targetting before suggesting something that changes the structure of execution.
At that point it's up to you (if you're using this interactively) or the coding agent you're doing to specify the broader constraints better.
does that answer your question?
Diagnosing performance bottlenecks take 4โ8x longer than fixing it.
We just shipped coding agent integrations for @nCompass_tech โ now Claude Code and Cursor can talk directly to our agent to diagnose performance bottlenecks for you.
Add it as an MCP. Keep coding like normal.
Docs โ https://t.co/GdYoTXqLo2
#CodingAgents #Vibetracing #CUDA #GPUKernels #ClaudeCode #Cursor
Last week we released AI agent integration for system level traces, i.e. nsys-reps and torch profile traces. Sharing a peak into what we're releasing soon for integrations with .ncu-reps for kernel developers!
Take a look here - https://t.co/mdQrmaHPMj for a demo of both and get started with agent assisted trace analysis.
For the first time you can run diffs and chat with all types of GPU performance related trace data all within your VSCode / Cursor IDE!
Excited to hear more around what parts of profiling and performance analysis you'd want to see more automation around. Reply in the comments below!
#MLSys #PerformanceOptimization #Profiling #nCompass
Got tired of manually reading through 200+MB trace files so we've now hooked them up to an AI agent and you can just chat with it. #vibetracing
Checkout our docs - https://t.co/GdYoTXqLo2
0:00 - Intro to the AI agent
1:00 - Example prompt with 2 questions
2:08 - Answer to question 1 - what should I care about in this trace
4:18 - Answer to question 2 - how do kernels map to operations from LLMs
Link to both the trace and responses in the comments!
#MLSys #GPUPerformance #PerfettoTraces #NsysReps
wrote a blog on how to use ncu directly on @vllm_project for a simple @Alibaba_Qwen model without taking 18+ hours to profile. I was tired of having to isolate kernels in a repro script to profile them with ncu. you can just isolate kernels direclty on a large trace and profile individual ones.
check it out here - https://t.co/fzJDABgTVA
#mlsys #vllm #inference #ai
we're introducing the concept of diffs to trace data - pick two traces, run the diff and get information on changes in durations, launch args, renames, additions and deletions of events in the trace.
check out how you can view diffs between @vllm_project 0.12.0 and 0.13.0 on the same workload - https://t.co/wjDtR8Dxhy
Ever wanted to run Nsight Compute (ncu) on a large codebase and understand individual GPU kernel performance โ without isolating each kernel?
Iโve hit this problem many times profiling systems like vLLM when asking the question - how well is a specific kernel performing for this workload?
In practice, that means using ncu. But with default settings, running ncu on thousands of kernel launches is prohibitively expensive โ so you end up building isolated repos just to profile one kernel.
The good news: you donโt have to.
Iโm writing a 3-part series on reducing ncu overhead enough to profile kernels directly in large, real-world codebases. Part 1 is out today.
Part 1 โ understanding the metrics ncu reports (important because reducing which metrics you profile is important to reduce overheads of profiling)
Part 2 โ how to actually apply this on large codebases
Part 3 โ why combining ncu + nsys beats using either alone
๐ Part 1: https://t.co/lsoUXKWpWy
If youโve run into this before, Iโd love to hear in the comments how you handle kernel-level profiling at scale.
#CUDA #GPU #PerformanceEngineering #NsightCompute #MLSys #NVIDIA #Profiling