Original Post
I'm getting so frustrated here at work. We're not doing anything graphics related anymore so no opportunity to expand my knowledge and there is no real resident graphics expert to learn from. So I'm reaching out to the experts on Gamedev to set me straight.
I want to get some experience/knowledge on optimizing shaders. I know that's a bit general, but I know its a big weak point of mine. Other than avoid expensive instructions, or dependent texture reads, I really don't have much of an idea on how to make any given shader run faster. I have some ideas, but honestly I have not done any profiling to see if any of them are true. Honestly I don't even know what tools to use so that I can profile. What do people use out there?
I would like to ask the experts out there...
1) Are dependent texture reads a real problem anymore on the newer cards? (I never really understood why they were so slow on the old cards other than maybe stalling the pipeline due to poor/non-existent thread scheduling??)
2) I remember reading somewhere that I should issue texture lookups asap so that there is work to do while we wait, is this true?
3) I also remember read about how the newer cards aren't really 4-way SIMD processors (as in operate on float4's) and that really they execute four instructions on four pieces of data in one clock cycle. Something about how NVIDIA analyzed shader usage and found that typical operations are not float4s but rather on floats. Does that mean I should not try and pack data into float4's in the shader? I'm not talking about packing data into a vertex attribute, I mean temp registers in the shader. Speaking of temp regs...
4) Now that everything is unified architecture I read that I should keep GPR usage to a minimum such that more threads can be executing in parallel. This seems to conflict with 3), although I would think 4) would outweigh 3).
5) I had a guys tell me once that every keyword in hlsl is takes 1 clock cycle (ex: dot( N, L) is 1 cycle, exp( X ) is 1 cycle, pow( X, Y), clamp(), max(), sin(), atan() etc). Is this true? He's a pretty senior guy but not really up on the new tech. I'm skeptical. I wouldn't even be so sure that one assembly instruction is equal to 1 cycle since different GPU's will decode the general assembly into something else, right? But I could be wrong (which is why I'm asking you)
6) I was in an interview once and a graphics guy told me that sometimes interpolators can be a bottleneck in the pipeline. I can understand interpolators (color) causing some precision issues, but speed!? How? I believe interpolators are separate hardware units so maybe they can't really handle a full load when all of the vertex attributes need interpolating. (just guessing)
I'm deeply interested in these kind of things to keep in mind when writing my shaders (or interviews). I understand some of these will depend on hardware, but I'm not content with that. I want to understand WHY something is the way it is. I would much rather hear something like "Well on the PS3 it will do X, on the 360 it will do Y, and on ATI it will do Z vs NVIDIA it will do W. Documentation or an article would be the best so that I can read about these things myself.
I want to get some experience/knowledge on optimizing shaders. I know that's a bit general, but I know its a big weak point of mine. Other than avoid expensive instructions, or dependent texture reads, I really don't have much of an idea on how to make any given shader run faster. I have some ideas, but honestly I have not done any profiling to see if any of them are true. Honestly I don't even know what tools to use so that I can profile. What do people use out there?
I would like to ask the experts out there...
1) Are dependent texture reads a real problem anymore on the newer cards? (I never really understood why they were so slow on the old cards other than maybe stalling the pipeline due to poor/non-existent thread scheduling??)
2) I remember reading somewhere that I should issue texture lookups asap so that there is work to do while we wait, is this true?
3) I also remember read about how the newer cards aren't really 4-way SIMD processors (as in operate on float4's) and that really they execute four instructions on four pieces of data in one clock cycle. Something about how NVIDIA analyzed shader usage and found that typical operations are not float4s but rather on floats. Does that mean I should not try and pack data into float4's in the shader? I'm not talking about packing data into a vertex attribute, I mean temp registers in the shader. Speaking of temp regs...
4) Now that everything is unified architecture I read that I should keep GPR usage to a minimum such that more threads can be executing in parallel. This seems to conflict with 3), although I would think 4) would outweigh 3).
5) I had a guys tell me once that every keyword in hlsl is takes 1 clock cycle (ex: dot( N, L) is 1 cycle, exp( X ) is 1 cycle, pow( X, Y), clamp(), max(), sin(), atan() etc). Is this true? He's a pretty senior guy but not really up on the new tech. I'm skeptical. I wouldn't even be so sure that one assembly instruction is equal to 1 cycle since different GPU's will decode the general assembly into something else, right? But I could be wrong (which is why I'm asking you)
6) I was in an interview once and a graphics guy told me that sometimes interpolators can be a bottleneck in the pipeline. I can understand interpolators (color) causing some precision issues, but speed!? How? I believe interpolators are separate hardware units so maybe they can't really handle a full load when all of the vertex attributes need interpolating. (just guessing)
I'm deeply interested in these kind of things to keep in mind when writing my shaders (or interviews). I understand some of these will depend on hardware, but I'm not content with that. I want to understand WHY something is the way it is. I would much rather hear something like "Well on the PS3 it will do X, on the 360 it will do Y, and on ATI it will do Z vs NVIDIA it will do W. Documentation or an article would be the best so that I can read about these things myself.