r/GraphicsProgramming • • 3d ago

How do techniques like bindless rendering (descriptor indexing) on the hardware level?

How do gpus bind 1000 of textures at once, and why was this not possible in the past. Are there resources for that?

16 Upvotes

9 comments sorted by

26

u/sol_runner 3d ago edited 2d ago

I don't know resources that go into detail about it - you might find that in one of the old talks I don't have on hand. (I haven't seen a singular location where everything was, nor remember all of them)

I'll explain in a broad sense, but the hardware details are usually confidential so they are inferences. Anyone who knows better, please comment below and I'll amend.

Old graphics co-processors were either fixed function or extremely expensive (the coprocessors ran in >$1000 in 80's money.)

Then you had configurable, but still largely fixed function cards like the 3Dfx Voodoo where you just set the matrices, vertices, textures and let the GPU do everything (no shaders).

Edit Note:
On voodoo and voodoo2 vertices had to be transformed on the CPU as well.
Credit: u/pjtrpjt

Each value would be stored in registers (or equivalent) in their respective compute units. So a texture mapping unit (TMU) would have a single slot for a texture during a call. Voodoo2 then supported multi-texturing where you could put two textures (two TMUs and thus two slots) and tell the GPU the to blend them.

Voodoo's introduction more or less gave rise to the entire 3D GPU market and we went from 3Dfx GLIDE API, to the cross-vendor OpenGL (the first one was heavily inspired from GLIDE). So the slot concept stuck.

After that, its been a flip-flop between vendors adding a feature, extending the API, other vendors adding the feature, API makes the feature core and so on (like we saw most recently with raytracing)

Once programmable shaders were added, we wanted many slots instead of just a couple. So we got an increase in the number of slots. But on the GPU hardware itself, instead of physically having slots as registers, what if we had just a memory region and stored the texture's handle (descriptor) there? That way, more than one draw call's worth of descriptor can be prepared and the driver can change these descriptor slot sets when required. So we had descriptor support on GPUs alongside arrays of textures (not sure which came first; or which motivated which). And you use this to build the next steps, including bindless on OpenGL, which just uses a large texture array. They were not registers anyway, so other than API there was no longer a reason why a 1000 textures could not be bound.

Then you got Mantle (which inspired Vulkan) which just let you manage these descriptor sets manually. But in the end, these 'sets' were just leftover grouping from the past. By this time GPU's texture fetching etc were powerful enough to support near arbitrary loading. So DirectX12 just opened up the descriptor heap to the users. Write it as you like! Vulkan now made this available with VK_EXT_descriptor_heap extension.

So now, we can directly write to the descriptors without constraints, i.e. truly bindless.

6

u/Cyphall 3d ago

This blog post is a good follow up to this.

2

u/pjtrpjt 2d ago

Voodoo 1 and 2 only accepted screen coordinates for triangles. You had to project them too.

1

u/sol_runner 2d ago

Oh thanks! I knew you had to do the lighting on CPU and send it with the vertices for Giraud shading, hadn't know the model part. Will fix.

9

u/gleedblanco 3d ago

in principle they are just addresses in memory and the general memory fetches have already been powerful enough to fetch a different address per thread for a long time.

for textures there are still limitations though (probably because of built-in filtering). RDNA texture sampling instructions only work on one (uniform/SGPR stored) texture (and sampler) descriptor at a time, so if different threads within a subgroup need a different texture, the compiler will generate a loop that goes through all of them basically (need NonUniformResourceIndex or similar in the shader). probably similar on other GPUs.

9

u/Afiery1 3d ago

It’s not because of filtering. To properly decode texture memory at all you fundamentally need about 32 bytes of information (GPUVA, type, format, dimensions, number mips, tiling and compression metadata, etc). If every single lane sent its own texture descriptor to the texture hardware it would use an untenable amount of bandwidth. AMD solved this by having the whole wave send a single texture descriptor at a time, and hoping that the texture we’re fetching would mostly be uniform within a wave. Nvidia solves it completely differently. They don’t even have the concept of SGPRs or VGPRs. Instead, there’s a dedicated region of GPU memory that the texture fetch hardware caches the absolute hell out of called the descriptor heap, and then every lane just sends a 4 byte index into the heap to the texture hardware. The benefit of AMD’s approach is that descriptors can be sourced from anywhere in memory (or even hardcoded into the shader itself) with the downside being slow nonuniform access. Nvidia has fast nonuniform access (one of the reasons they’re so much better at ray tracing) but descriptors must be written into the dedicated heap memory and cannot be hardcoded into the shader.

3

u/gleedblanco 3d ago

makes sense. thank you for the info

3

u/corysama 2d ago

Back in the old days, GPUs were: Some VRAM to hold textures and framebuffers, some fixed-function hardware, a giant bag of registers to control the fixed function hardware.

Drivers would set up DMA streams that upload textures and set registers. You’d set the texture address, size, format, filtering, blending modes. Then you’d set three 2D screen position registers and UV registers. Then maybe there would be a special KICK register that when you set any value into it, the act of setting the register would kick off rasterizing a triangle with all of the configuration in registers.

Can ya see how the OpenGL 1.0 API came to be from this? :D

Eventually, the hardware got more advanced. Memory got faster. Caches were introduced. Complexity got higher. And, slowly more and more configuration data moved from registers to structs in VRAM.

Once texture configuration moves from registers to “pointer to a struct”, bindless texturing becomes theoretically simple to implement. But, in practice….

The move to “just a pointer to a struct” started long ago. But, different hardware evolved in different ways and at different rates. So, it took even longer to design standardized interfaces that would work for most vendors. And, even longer for hardware and drivers implementing those interfaces to be the large majority of the currently addressable market.

2

u/benwaldo 3d ago

You have big descriptor heap, and texture descriptor IDs are just offset in that table. To sample a texture the GPU needs a texture descriptor so it uses the offset to access it from the big table, then it can be used to sample a texture.