ANE INT8 on A20 Pro is about 105 TOPS for INT8. Measured with my little MPSGraph based app, https://t.co/f2EbaZGlkI. Note that these numbers are likely Winograd 1D F(2, 3) "accelerated".
@anemll Mostly NO. From .hwx and reverse engineering of the ANECompiler, current 1D Winograd implementation doesn’t support fp8. If Winograd is enabled, you should be able to find it in .hwx file.
@anemll So far it’s for A20 Pro (h19) and M6 (h18g) only from the ANECompiler in 27.0 RC. I guess although M5 Ultra has 2 ANE units, there is fast enough bus/interconnect between them.
@anemll Assuming the question is: does the ANECompiler handle cases with > 2 ANE units? Nope, not for the bonded network of the ANECompiler, which is hardcoded to support exactly 2 units in macOS/iOS 27.0 RC. BTW, I have a short writeup on the bonded network at https://t.co/nNilmHs4dT
Dug into ANECompiler.framework of iOS/macOS 27 RC and found “bonded networks” on H18g/M6 and H19/A20 Pro. Before, multi-ANE chips like the M-Ultra just split up the work at the driver level across separate devices. (1/2)
Dug into ANECompiler.framework of iOS/macOS 27 RC and found “bonded networks” on H18g/M6 and H19/A20 Pro. Before, multi-ANE chips like the M-Ultra just split up the work at the driver level across separate devices. (1/2)
But with bonded networks, the compiler actually splits a single model across two physical ANE units itself. So, I guess there is relatively high-speed interconnect between the two ANE units. The compiler spits out a .hwx file with both bonded and non-bonded versions. (2/2)
@anemll fp8 and fp4 are in the MLIR path already. As far as I can rember, they are in ANECompiler, too. I didn't check the CoreML and MIL part. I mean it should be possible to use FP8 with CoreAI or MPSGraph.
Finally figured out how to read Apple ANE's PMU counters. In short, you have to change boot-args. I put down some example cdoe at https://t.co/25Fmd7L6uH. Tested on M1 and M4 only.
@anemll@originalmaderix I know why. I measure ANE compute capacities couple months ago. AFAICR, MIL doesn't support INT8. With MPSGraph, which uses MLIR, you can get expected INT8 performance.
https://t.co/L3gNigJdHA
MLCommons® just launched MLPerf® Mobile on the Google Play Store! 📱
Benchmark your Android device’s AI performance on real-world ML tasks with this free, open-source app.
Try it now: https://t.co/YKN4OtzU1k
Just built an MCP for Ghidra.
Now basically any LLM (Claude, Gemini, local...) can Reverse Engineer malware for you. With the right prompting, it automates a *ton* of tedious tasks.
One-shot markups of entire binaries with just a click.
Open source, on Github now.