Skip to main content
GameDev.net gamedev.net
🔒 Locked

Shader doesn't work on AMD GPUs with latest driver

Started by Beosar Apr 4, 2019 at 3:01 PM 17 replies 4.9k views
Original Post
Beosar
Beosar

Hi,

I use the following shader (HLSL) to scale positions, normals, and depth for post processing:



Texture2D normalTexture : register(t0);
Texture2D positionTexture : register(t1);
Texture2D<float> depthTexture : register(t2);


static const uint POST_PROCESSING_SCALING_MASK = 0xC0;
static const uint POST_PROCESSING_SCALING_FULL = 0x00;
static const uint POST_PROCESSING_SCALING_HALF = 0x40;
static const uint POST_PROCESSING_SCALING_QUARTER = 0x80;


cbuffer PixelBuffer{
    float fogStart;
    float fogMax;
    uint Flags;
    float Gamma;


    float3 fogColor;
    float SpaceAlpha;


    float FrameTime;
    float MinHDRBrightness;
    float MaxHDRBrightness;
    float ScreenWidth;


    float3 SemiTransparentLightColor;
    float ScreenHeight;


    float SpaceAlphaFarAway;
    float SpaceAlphaNearPlanets;
    float HDRFalloffFactor;
    float InverseGamma;


    float3 CameraLiquidColor;
    float CameraLiquidVisualRange;


    float CameraInLiquidMinLerpFactor;
    float CameraInLiquidMaxLerpFactor;
    float MinCloudBrightness;
    float MaxCloudBrightness;


    float4 BorderColor;


    float3 SunColor;
    float padding0;
};


struct PixelInputType{
    float4 position : SV_POSITION;
};



struct PixelOutputType{
    float4 normal : SV_Target0;
    float4 position : SV_Target1;
    float depth : SV_Depth;
};


PixelOutputType main(PixelInputType input){
    PixelOutputType output;
    const uint ScalingFlag = (Flags & POST_PROCESSING_SCALING_MASK) >> 6;


    uint3 TexCoords = uint3(uint2(input.position.xy) << ScalingFlag, 0);
    const uint Max = (1 << ScalingFlag) << ScalingFlag;
    uint UsedIndex = 0xFFFFFFFF;
    for (uint i = 0; i < Max && i < 16; ++i) {
        const uint3 CurrentTexCoords = TexCoords + uint3(i & ((1 << ScalingFlag) - 1), i >> ScalingFlag, 0);
        output.position = positionTexture.Load(CurrentTexCoords);
        if (output.position.w >= 0.0f) {    
            UsedIndex = i;
            break;            
        }        
    }
    if (UsedIndex == 0xFFFFFFFF) {
        output.normal = float4(0, 0, 0, 0);
        output.depth = 1.0f;
        discard;
    }
    else {
        const uint3 CurrentTexCoords2 = TexCoords + uint3(UsedIndex & ((1 << ScalingFlag) - 1), UsedIndex >> ScalingFlag, 0);
        output.normal = normalTexture.Load(CurrentTexCoords2);
        output.depth = depthTexture.Load(CurrentTexCoords2);
    }
    return output;
}


It doesn't work with the latest drivers on AMD graphics cards, neither on my laptop (R7 M270) nor one of my friend's computers (RX 480).

It works perfectly on Nvidia/Intel GPUs and it worked with older drivers on my laptop, too.

On AMD GPUs, it outputs wrong data, i.e. output.position.w < 0 even though that is impossible (if it is < 0, UsedIndex will be 0xFFFFFFFF, therefore the pixel will be discarded). If I manually set output.position.w = 0 in both of the branches on (UsedIndex == 0xFFFFFFFF), it outputs 0 in the w component. If I set it in only one of the branches (doesn't matter which one), then output.position.w is negative in the output.

The other output (positions.xyz, normals.xyz, and depth) seems to be correct, but I'm not sure about normals.w.

It literally makes no sense and I think it's a driver bug, what can I do?

Cheers,

Magogan

Edit: If I add the following code before returning output, it gets even weirder:


if (output.position.w < 0) {
	output.position.w = 0;
}
else {
	output.position.w = 1;	
}

After that, output.position.w is neither 0 nor 1, but something such that it is not >= 0, so probably NAN or negative. WTF?

JohnnyCode
JohnnyCode

How do you declare your GBuffer multiple render target buffers?

If shader is compiled and linked, it should not be the source of problem at all.

Beosar
Beosar
1 minute ago, JohnnyCode said:

How do you declare your GBuffer multiple render target buffers?

What?

1 minute ago, JohnnyCode said:

If shader is compiled and linked, it should not be the source of problem at all.

You would think so... And you would be wrong...

JohnnyCode
JohnnyCode

My question was quite clear, your pixel function outputs to 3 render targets of likely float components per channels.

You also sample one float per component texture.

If you think modifying your shader will solve your issue while it was succefully compiled, good luck then.

Beosar
Beosar
Just now, JohnnyCode said:

My question was quite clear, your pixel function outputs to 3 render targets of likely float components per channels.

Yes? It just works like that. I don't understand the question.

joeblack
joeblack

if its driver version related, it could be driver bug, you can write to amd support about it.

Beosar
Beosar
Just now, joeblack said:

if its driver version related, it could be driver bug.

I'm like 80% sure it is, but how do I get AMD to fix it, without my game being as popular as The Witcher or Battlefield? How can I find a workaround in the meantime?

joeblack
joeblack

you can write to amd forums/support mail. it looks that problem is in positiontexture. Have you checked content of that texture ?

Beosar
Beosar
2 minutes ago, joeblack said:

you can write to amd forums/support mail.

And you think they will fix it?

2 minutes ago, joeblack said:

it looks that problem is in positiontexture. Have you checked content of that texture ?

Yes, the texture is correct. Otherwise it wouldn't work with Nvidia/Intel GPUs (and an older version of the AMD driver on my laptop). And even if the texture is not correct, that wouldn't explain the behavior I observed.

joeblack
joeblack

Well it depends, on AMD speed.

What is format of texture ? did you checked shader that is writing to it ?

Beosar
Beosar
11 minutes ago, joeblack said:

What is format of texture ?

R32G32B32A32_FLOAT

11 minutes ago, joeblack said:

did you checked shader that is writing to it ?

They are correct. It works without scaling, i.e. if I don't use this shader and just use the full resolution position texture directly.



... 
  
SamplerState Sampler
{
	Filter = MIN_MAG_MIP_POINT;
	AddressU = Clamp;
	AddressV = Clamp;
};


...
  
	float2 TextureSize = float2(0, 0);
	positionTexture.GetDimensions(TextureSize.x, TextureSize.y);
	const uint Max = (1 << ScalingFlag) << ScalingFlag;
	uint UsedIndex = 0xFFFFFFFF;
	for (uint i = 0; i < Max && i < 16; ++i) {
		const uint3 CurrentTexCoords = TexCoords + uint3(i & ((1 << ScalingFlag) - 1), i >> ScalingFlag, 0);

		output.position = positionTexture.Sample(Sampler, (float2(CurrentTexCoords.xy)+0.5f)/ TextureSize);
		if (output.position.w >= 0.0f) {	
			UsedIndex = i;
			break;			
		}		
	}
	
	
...

With this change it works?!?


Edit: Adding [unroll] to the loop also fixes the problem. But why?

joeblack
joeblack

Well, i think that amd should answer that. at least you have workaround

Beosar
Beosar
1 hour ago, joeblack said:

at least you have workaround

Which costs like 10-20% performance (FPS)...

JoeJ
JoeJ

AMD has already fixed 2 driver bugs i had submitted in their forum. Usually they want a repo of your code so they can reproduce.

Beosar
Beosar
5 minutes ago, JoeJ said:

in their forum

Which forum? I've only found the DevGurus forums, are those the correct forums for that?

I'm not sure if I should give them access to my complete game code...

JoeJ
JoeJ
28 minutes ago, Magogan said:

DevGurus forums

Yes

28 minutes ago, Magogan said:

I'm not sure if I should give them access to my complete game code...

Yeah... unfortunately it is some work to strip it down so only the necessary things to show the bug remain. :|

JohnnyCode
JohnnyCode

Usualy it is unproper intitaion of texture/render target buffers, that some drivers omit and forgive, some do not. That is why I asked for how you establish your multiple render targets, and textures, how you set them for pipeline, do you check for available frame bufffer formats on device initiation etc.?

Vilem Otte
Vilem Otte
On 4/4/2019 at 8:44 PM, JoeJ said:

Yeah... unfortunately it is some work to strip it down so only the necessary things to show the bug remain. :|

Ideally you need to have the smallest source possible that reproduces the bug, pushing them whole code base is not going to work well. Also make sure to give them which version of driver was working, and which one did break it.

20 hours ago, JohnnyCode said:

Usualy it is unproper intitaion of texture/render target buffers, that some drivers omit and forgive, some do not.

From experience (and especially with Intel gpus and my work on our custom modified GTK+ where we use lots of OpenGL) - make sure to not just read articles but exact definitions of functions from the standard. NVidia tends to be the least forgiving, while AMD and Intel gpus tend to be much less forgiving and will give you an error each single time.


To the problem:

On 4/4/2019 at 6:40 PM, Magogan said:

Edit: Adding [unroll] to the loop also fixes the problem. But why?

There is major problem with loops in shader source code, as every shader compiler tries to optimize it out. When working with loops - it needs to determine whether you want to implement loop as actual loop (with condition and jump instructions - which is a standard way), or unroll the loop (and use multiple jumps just to break after i-th iteration if condition is satisfied).

In case of some loops the optimization may simple fail and give you incorrect results. If possible I recommend you give additional parameter do define whether you want to [unroll], [loop], [fastopt] and then there is another attribute for compute shaders [allow_uav_condition] (specifying [unroll] means that all other are ignored, also [loop] and [unroll] are mutually exclusive).

As you have 2 conditions in the loop, I assume that compiler simply failed to optimize your loop and produced invalid code. It is possible you received warning about it in the shader compilation output.


Can you try the same shader with [loop] attribute only?

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.