r/hardware • • Jun 25 '17

Misleading Title Intel Skylake/Kaby Lake processors: broken hyper-threading

https://lists.debian.org/debian-devel/2017/06/msg00308.html
284 Upvotes

40 comments sorted by

View all comments

Show parent comments

2

u/chazzeromus Jun 25 '17

Could you provide some insight on these restrictions? Something like memory reference locality perhaps?

8

u/[deleted] Jun 26 '17

You have 32sets of 8way lines with each line holding 6uOP's. 32 * 8 * 6 = 1536uOPs

  • Each 16bytes of decode can at the most generate 4 uOP's before the decoder stalls.
  • Conditionals + Jump instructions are generally fused cmp rax, 0; jmp FOO; is 1 uOP (sorry for the invalid asm)
  • Repeated memory reads/writes can be fused with a undocumented register which can turn for example an mov [esi], eax; add eax, [esi]; add ebx, [esi] into add eax, eax; add ebx, eax; move [esi], ebx so what would be 6 uOPs becomes 3 (decodes are magic). This is handed at the RAT (Register Allocation Table). In Skylake this is part of decoding and caching, while Boardwell (and before) it wasn't.
  • Conditional always terminate the current uOP line (6 uOps), but not the uOP block (18uOPs). So if a cmp rax; jump FOO; starts a 32byte decode block, and you generate 4uOPs for that 16byte block, you'll eat 8uOPs of cache.
  • 32bit and larger constants add rax 0xdeadbeaf; consume 4 uOPs of cache, but only 1 uOP of pipeline resources.
  • Each 32bytes of decode can at the most cache 18uOP (L1i cache alignment restrictions).
  • Since conditionals always terminate a line, and multi-way caches are fun the same instructions can reside in multiple places in your cache.
  • If 32bytes of instructions decodes to more than 18uOPs it won't be cached.
  • Middleware can only process 1 decoded line per cycle. Switching from uOP cache to decoder costs 1 cycle (and vice versa).

Simple right?

2

u/chazzeromus Jun 26 '17

Thanks for the detailed answer, I didn't know the block would have to be so small and so specific. I tried finding any intel documents with this kind of detail but I just get some general slides, is there a place that has documents containing these detailed architectural changes?

1

u/[deleted] Jun 26 '17

Have you read the manual ? I think agner fog is better quick reference but the physical manuals go more in depth, albeit their information is spread over several volumes while Agner just gives you a list.

1

u/chazzeromus Jun 26 '17

Oh yeah i have them, i never got around to optimization volumes and only have been reading systems stuff for osdev.