@nCompass_tech is a performance optimization dev tool that unifies ๐ฝ๐ฟ๐ผ๐ณ๐ถ๐น๐ถ๐ป๐ด, ๐๐ฟ๐ฎ๐ฐ๐ฒ ๐ฐ๐ผ๐น๐น๐ฎ๐ฏ๐ผ๐ฟ๐ฎ๐๐ถ๐ผ๐ป and ๐๐ฟ๐ฎ๐ฐ๐ฒ ๐ฎ๐ป๐ฎ๐น๐๐๐ถ๐ of AI systems.
We automate the process of identifying and fixing the root cause of performance bottlenecks across all levels of the AI infrastructure stack. Our agent automatically surfaces key bottlenecks and works in tandem with Claude Code / Cursor to implement the fixes.
Using our tool, we were able to automatically improve a Hopper GEMM kernel taking it from ๐ฏ๐ฌ% ๐๐น๐ผ๐๐ฒ๐ฟ ๐๐ผ ๐ฏ% ๐ณ๐ฎ๐๐๐ฒ๐ฟ than a near optimal CUTLASS kernel within a day - this took us months before!
Weโre excited to officially release v0.1.0 of our VSCode extension that includes ๐ป๐ฎ๐๐ถ๐๐ฒ ๐๐ฟ๐ฎ๐ฐ๐ฒ ๐๐ถ๐ฒ๐๐ถ๐ป๐ด, ๐๐ฟ๐ฎ๐ฐ๐ฒ ๐ฑ๐ถ๐ณ๐ณ๐, ๐๐ฟ๐ฎ๐ฐ๐ฒ ๐ฐ๐ผ๐น๐น๐ฎ๐ฏ๐ผ๐ฟ๐ฎ๐๐ถ๐ผ๐ป and integration of our ๐๐ ๐ฎ๐ด๐ฒ๐ป๐ directly into Cursor / Claude Code. Weโre working closely with an initial set of users including the core maintainers of @vllm_project to build features that best support the community.
๐๐ณ ๐๐ผ๐โ๐ฟ๐ฒ ๐ฎ๐ป ๐ฒ๐ ๐ฝ๐ฒ๐ฟ๐ ๐๐ฃ๐จ ๐ฝ๐ฒ๐ฟ๐ณ๐ผ๐ฟ๐บ๐ฎ๐ป๐ฐ๐ฒ ๐ผ๐ฝ๐๐ถ๐บ๐ถ๐๐ฎ๐๐ถ๐ผ๐ป ๐ฒ๐ป๐ด๐ถ๐ป๐ฒ๐ฒ๐ฟ - bring your time spent optimizing performance down to days rather than weeks - view, collaborate and analyze faster than you do today.
๐๐ณ ๐๐ผ๐โ๐ฟ๐ฒ ๐ป๐ฒ๐ ๐ฎ๐ป๐ฑ ๐๐ฎ๐ป๐ ๐๐ผ ๐น๐ฒ๐ฎ๐ฟ๐ป ๐ต๐ผ๐ ๐๐ผ ๐ผ๐ฝ๐๐ถ๐บ๐ถ๐๐ฒ ๐ฝ๐ฒ๐ฟ๐ณ๐ผ๐ฟ๐บ๐ฎ๐ป๐ฐ๐ฒ - you now have an expert pair-programmer to consult at any time.
๐๐ฒ๐ ๐๐๐ฎ๐ฟ๐๐ฒ๐ฑ ๐ณ๐ผ๐ฟ ๐ณ๐ฟ๐ฒ๐ฒ ๐๐ผ๐ฑ๐ฎ๐ using VSCode, Cursor or Claude Code - https://t.co/lWXuYeoavi
We also work directly with you to identify opportunities for performance improvements. If you would like us to ๐ฎ๐๐ฑ๐ถ๐ ๐๐ผ๐๐ฟ ๐๐๐ฎ๐ฐ๐ธ and suggest ways to improve performance - reach out at [email protected]!
Got tired of manually reading through 200+MB trace files so we've now hooked them up to an AI agent and you can just chat with it. #vibetracing
Checkout our docs - https://t.co/GdYoTXqLo2
0:00 - Intro to the AI agent
1:00 - Example prompt with 2 questions
2:08 - Answer to question 1 - what should I care about in this trace
4:18 - Answer to question 2 - how do kernels map to operations from LLMs
Link to both the trace and responses in the comments!
#MLSys #GPUPerformance #PerfettoTraces #NsysReps
wrote a blog on how to use ncu directly on @vllm_project for a simple @Alibaba_Qwen model without taking 18+ hours to profile. I was tired of having to isolate kernels in a repro script to profile them with ncu. you can just isolate kernels direclty on a large trace and profile individual ones.
check it out here - https://t.co/fzJDABgTVA
#mlsys #vllm #inference #ai
Ever wanted to run Nsight Compute (ncu) on a large codebase and understand individual GPU kernel performance โ without isolating each kernel?
Iโve hit this problem many times profiling systems like vLLM when asking the question - how well is a specific kernel performing for this workload?
In practice, that means using ncu. But with default settings, running ncu on thousands of kernel launches is prohibitively expensive โ so you end up building isolated repos just to profile one kernel.
The good news: you donโt have to.
Iโm writing a 3-part series on reducing ncu overhead enough to profile kernels directly in large, real-world codebases. Part 1 is out today.
Part 1 โ understanding the metrics ncu reports (important because reducing which metrics you profile is important to reduce overheads of profiling)
Part 2 โ how to actually apply this on large codebases
Part 3 โ why combining ncu + nsys beats using either alone
๐ Part 1: https://t.co/lsoUXKWpWy
If youโve run into this before, Iโd love to hear in the comments how you handle kernel-level profiling at scale.
#CUDA #GPU #PerformanceEngineering #NsightCompute #MLSys #NVIDIA #Profiling
You wouldnโt collaborate on code in Google Docs โ so why do we still do that with trace files?
We just shipped a new feature in the nCompass VSCode / Cursor extension that makes trace collaboration actually usable:
โข Save traces to secure cloud storage
โข Share traces via a link teammates can open directly in their IDE
Our team is partially remote and works with a lot of traces. Until now, sharing meant SCP or Slack uploads. Now we can share traces straight from remote servers without leaving the IDE.
As teams push harder on AI inference and training performance, collaborating on traces is becoming as important as collaborating on code. Weโre building the tooling for that workflow.
Tutorial: https://t.co/00lNjLzklM
#MLSys #Profiling #GPU #VSCode #Performance
๐ Shipped nCompass SDK v0.1.12 - performance debugging shouldnโt require permanent code changes.
v0.1.12 makes it effortless to inject NVTX / TorchRecord markers at runtime โ without editing your codebase or maintaining a debug branch.
How it works:
โข Select a region in VSCode / Cursor
โข pip install ncompass>=0.1.12
โข Set env vars + run with torch profiler / nsys
โข Markers show up in the trace โจ
Fully open source. New injectors welcome.
๐ฅ Tutorial: https://t.co/iWbBvtOMOs
โก Quick start: https://t.co/4oeilSzXgt
#MLSys #Profiling #GPU #Performance #PyTorch
As the saying goes, an image is worth a thousand tokens. So, to see a whole new line of @Alibaba_Qwen models released without multimodality was rather surprising. This thread is my breakdown of the state of multimodality in open and closed source AI.
Portkey now supports nCompass integration!
@nCompass_tech is known for its blazing-fast, serverless LLM deployments with no rate limits and flexible on-prem options.
With this integration, you can now:
โ Route and monitor LLM traffic to nCompass
โ Apply guardrails and track prompt-level metadata for every nCompass call
โ Track 40+ key metrics including cost, tokens, latency, and more
Read more about the integration here - https://t.co/olDaOim6gc
๐งญ@nCompass_tech (YC W24) is a platform for simplified hosting & acceleration of open-source LLMs. Its API requires only one line of code to integrate low-latency versions of these models into your AI pipeline.
https://t.co/utU9YPURNm
Congrats on the launch @adityaraja0 & @DiederikVink!