Seine Mutter wollte, dass er orthodoxer Priester wird. Stattdessen wurde er Bankräuber und Herrscher über ganz Osteuropa. Er brachte Russland die schnelle Industrialisierung - obwohl er gar kein Russe war - und tötete Millionen Russen. Heute kann man froh sein, wenn deutsche Jugendliche ihn wenigstens auf Erwachsenenbildern erkennen.
How aware are modern compilers of exact microarchitectural layouts?
Quite a lot…in one very specific way.
Intel x86 is *not* the same as AMD x86.
Sure…it’s the “same ISA” in the broadest sense, but the individual instructions often take different numbers of cycles.
You want to stall as little as possible, so the order your instructions are arranged is somewhat important. Technically, something like “AMD Zen 3” ordering on an Intel Skylake is sub-optimal.
If you look at LLVM, you’ll notice these X86Sched*.td (TableGen) files. Just about every x86 CPU generation has their own version.
It defines things like ROB Size, issue width, and misprediction penalties (in terms of cycles). What’s fascinating is how much of this is “guessed” (reverse-engineered) vs “revealed” by the hardware vendors.
From what I understand, Intel/AMD/etc will *sometimes* lend a helping hand / give some hints…but less than you’d expect.
It’s a very weird situation when you think about it. I assume vendors are tight-lipped about exact latencies for competitive secrecy…yet those are the exact things you’d need to know to extract the most performance out of your compiler!
If anyone knows other reasons for the secrecy, I’d love to hear it!
„Doch was ist unsere Ideologie wert, wenn sie nicht auch angewandt wird? Der Markt reagiert nicht wie ein kleines Kind, das beim ersten Anzeichen von Unwohlsein gleich zu schreien anfängt. Der Markt ist geduldig und wir müssen es auch sein.“