Obsessively exploring simulations, graphics, machine learning & AI. This is my demo space & devlog for side projects in WebGPU, WebGL, Metal & MLX/CUDA.
1st real milestone of realtime threejs meshlets + path-traced lighting in the browser. V2 (soon) will use shareable assets, so I'll have a live demo then. Lots to fix: pop-in, blotchy light, fireflies, etc. Opus 5.5 got me over the finish line. #threejs#webgpu#pathtracing
@DennisSmolek Hope you do btw, I feel like OIDN has a lot of room to squeeze. Good luck! Sorry for the crazy response, ONNX is like one of my pet peeves haha.
Neural denoiser v0.4 is trained + multiple spp frame support lands milestone 2 for my #threejs real-time pathtracer. This version revisits regular meshes. Normal scenes work w/ many moving lights, glass, caustics, mats, skinned mesh, morphs, & GI. Captured w/ settings on 'ultra'.
@DennisSmolek I was literally only building this so that my webgpu agent swarm could have a nicer 3D home to navigate when showcasing its shared decision making and self weight mods through accelerated backprop.
@DennisSmolek I actually have a whole AI runtime I'm releasing. Like 95% done. That's my main focus but this was a side quest. ONNX has really limited use cases in my opinion. I highly recommend dropping it beyond PoC. I have a hard time imagining anything being better for having used it.
@garrettkjohnson The meshlet pathtracer is really about handling big scenes over being accurate. Closer to a UE approach. It can REALLY handle some big scenes. Hopefully releasing demos this week that might surprise folks. I know it shocked me even when it ran poorly on my iPhone 13.
@garrettkjohnson The meshlet pathtracer is wavefront with streaming cluster LOD traced through the raster. Materials graphs & real-time complexity. 2 DAG & multi-level BVH w/ CLAS like structure. Trace is LOD aware & both versions work tightly w/ neural denoiser. Will converge both into one soon.
@garrettkjohnson My normal pathtracer is a megakernel with hybrid techniques... radiance cache, temporal, neural denoiser & sitting on a custom GPU built and resident BVH with full (but inaccurate) light transport. Big scenes get hurt but optimizing... Fixed frame costs are a double edge sword.
@garrettkjohnson Also there are three things happening here. The meshlet pathtracer, my normal pathtracer and three-gpu-pathtracer. Yours is wavefront sample accumulation over a queue focused on ground truth. Unbiased, good convergence. Actually correct. I assume by a mile.
@garrettkjohnson So I only knew your webgl version before & only glanced your webgpu version at release so I might be incorrect with this... but there are some pretty large architectural differences around tradeoff and intent.
@CodyBennett@jonathankennell By the way, not sure how your 2nd DAG is going but I suspect I'm going to deprecate in favor of more dedicated ray-only reps unless I can find a case where it really makes sense to keep for me.
Quick Correction - Video says meshlet pathtracer. This is not using meshlets, just left the title in by mistake. Two separate renderers but shared methods in many cases hopefully will have them combined at the end.
@mogmek Neural denoiser is overcorrecting right now and blotting out specular so I needed to combine my classic denoiser with the other for this capture.
@mogmek This is ramped up settings for sure... captured from my M5 pro on ultra settings. Not the earlier M1 Pro capture I was using for previous posts. But honestly, once I fix a few things on the neural denoiser this should be pretty close on the M1 as well at 30fps.
@CodyBennett@jonathankennell Also yes sorry to answer the implied question... two DAG, roughly same split and BLAS/TLAS still in there just avoiding when possible. Cache is already frail.
@CodyBennett@jonathankennell yeah, I've been following your efforts. I avoided BLAS/TLAS (or rather tried it & agreed the refit is faster in many cases). I've got a similar hybrid approaches but mostly settled on a weaker version of NVIDIA's CLAS.
Also I find it handy to set up these scripts in a little tailscale synced socket server so AI can debug issues simultaneously across devices. Just a nice little workflow I found for me at least.