Skip to main content
GameDev.net gamedev.net
🔒 Locked

First person weapon with different FOV in deferred engine

Started by Lutz Jul 6, 2011 at 6:45 AM 7 replies 6.8k views
Original Post
Lutz
Lutz
Hey,

I was wondering what people do for the first person weapon in a deferred engine. I know that some engines render the first person weapon with a different field-of-view, so the artists can make it look exactly the way they want. Now I was thinking of doing the same thing for a deferred engine. In the lighting pass, I have to get back from pixel coordinates to view space. Now the view space is different for normal pixels vs FP weapon pixels. So unless I store the view information in the g-buffer as well, it should be technically impossible to calculate the correct pixel position, right? What do other deferred engines do? Hybrid approach (render weapon with forward renderer and everything else deferred)?

- Lutz
Danny02
Danny02
I would just calc the lighting with the normal view matrix(which was used for the world), because otherwise u would get strange lighting. Like if you photoshop a pictureof a gun into a scene but the gun was lightened from a different angle, everything would be correct but it would look strange.
Hodgman
Hodgman
I've not thought through the implications of this suggestion (so exercise caution), but --

You could transform your vertices with BOTH the "real" view-proj matrix, AND with the artist's fudged view-proj matrix.
Give the 'fudged' positions to the hardware (i.e. in the POSITION semantic), but also pass through the 'real' positions to the fragment shader, and write the 'real' values into the g-buffer.
Lutz
Lutz
@Danny02: The problem is in a deferred engine you don't do the lighting per object, but write lighting information to the g-buffer and do the lighting afterwards in a pixelshader. At that stage, I don't know anymore which pixel was rendered with which view matrix.

@Hodgman: That's an interesting idea, but unfortunately I don't store xyz in the g-buffer, but only z and then I reconstruct xy from z and the view-proj matrix.
Danny02
Danny02
yeah I know, what I wanted to suggest was that u just ignore the fact that u rendered the weapon with an other fov and use ur normal view matrix for the hole picture, because if u would use somehow 1 view matrix for the scene and another for the weapon u would get strange lighting.

Besides, if u want to test the difference, u can just use the stencil buffer(would have to do the lighting 2 times(but shouldn't be that much of a problem.
MJP
MJP
You could override the depth output in your pixel shader, which would let you accomplish what Hodgman suggests. This would kill early z-cull, but that may not be a big deal since you're pretty much guaranteed that most of the pixels are going to pass the Z-test anyway.
Zoner
Zoner
We have done this with a forward renderer for Brothers in Arms:Hell's Highway and Borderlands. Aliens:Colonial Marines is deferred and uses nearly the same setup for the weapons. I've always called it the foreground FOV hack, it is becoming less hack like and more like a real feature But it still a hack because the render pipelines on the platforms we develop for are fairly different from each other, and sometimes needs some hand holding when the pipeline is modified. We work pretty much entirely with DX9 so some approaches can be improved by moving to 10/11 (particularly in the case of sampling the depth buffer as a texture).

First off, the good news:

Provided the near and far planes are the same for the projections, the depth values will be equivalent. What is different is the FOV is different so the screen space XY positions of the pixels will come out different, which generally only matters when de-projecting a screen pixel back into the world. For most effects that do this, the depth is usable as-is for depth based fog, and whatnot. This only requires making your artists cope with the same near plane as the world (and they will beg and scream for a closer plane for scopes and things that get right up on the near plane, but you have to say NO!, you get the custom FOV but you get it with this limitation).

And the bad news:

The depths are quite literally equivalent, which means the weapon will draw 'in the world' and have the rather annoying behavior of poking through walls you walk up to. So the fix is to render the gun several times, primarily for depth-only or stencil rendering. One possibility (and the one we use) is to render the gun with a viewport set to a range of 0/0.0000001 in order to get the gun to occlude the world. This is good for performance reasons, but bad if you have post process effects that absolutely must be able to to sample pixels from 'behind the gun'. This is a trade off someone has to sign off on. Performance usually wins that argument though, so we have opted to have the guns occlude everything (including hardware occlusion queries!). Another possibility is to render a pass to create a stencil mask of the weapon and occlude with that, but there are some complications that need to be understood, which I will talk about down near the bottom of this post.

Forward renderers can just draw the gun later in the frame at their leisure for the most part, after clearing depth (another thing I need to explain later) and drawing the gun. Deferred rendering doesn't have it as easy, as you need the gun to exist properly in the GBuffers when doing lighting passes, for both performance reasons, and to accept and real-time shadowmaps properly along with the world.

More good news:

Aside from the case of the gun or the player needing to cast shaders, the weapon will generally look just fine lit (and shadowed!) with the not-quite correct de-projected world space position. The depth will be completely correct from the view-origin's point of view, and the gun itself wont be too far from a correct XY screen position so it will just light and shadow just fine. UNLESS you attach a tiny light directly to the gun, at which point the light needs its position adjusted to be in the guns coordinate system instead of the world, so the gun looks correct when lit by the light. Muzzle flash sprites and whatnot have a similar problem, but in reverse, in that the sprite needs to be placed in the world correctly relative to the gun's barrel.

More bad news:

Getting the gun into the GBuffer properly and without breaking performance can be a bit tricky. We store a version of the scene depth's W coordinate in the alpha channel of the framebuffer (which is also the same buffer that is the the GBuffer's emissive buffer when it is generated). This is true of the PC and PS3. Rendering is basically Clear Depth, Render Gun Depth to the super-tight viewport, Render Depth, Render Scene GBuffer, Clear Z, Render Gun to GBuffer, perform lighting, translucency, post process etc. We can clear Z in this setup because the rest of the engine reads the alpha channel version of the depth for everything. The XBOX version read's the Z-buffer directly, so we have to preserve world depth values, so instead of 'Clear Z' we render the GUN twice, once a depth-always write and the second with the traditional less-equal test. This is necessary because the viewport clamped depths are not something you want the game to be using. This particular method is an extremely bad idea for PC DX9 hardware in general (NVIDIA's in particular).



The hardware is going to fight you:

You might be tempted to use oDepth in a shader. This is a bad idea, in that it disables the early depth & stencil reject feature of the hardware when the pixel shader outputs a custom depth. It is also not necessary for getting guns showing up correctly with a custom FOV. It is also a bad idea because you will also need to run a pixel shader when doing depth-only rendering, and it is extremely slow to do this (hardware LOVES rendering depth-only no-pixel-shader setups!). This is also the same reason why you should limit allowing masked textures to be used in shadowmap rendering, as they are significantly slower to render into the shadowmap (somewhere between 4 and 20x slower, its kind of insane how big of a difference it can be).

Getting it visually correct is not the real challenge. The real challenges lie in how many ways the hardware can break and performance can go off a cliff.

The early-depth and early-stencil reject behaviors of the hardware are particularlly finiky, which gets progressively worse the older the hardware is. NVIDIA's name for these culls are called ZCull and SCull. ATI(AMD whatever) calls it Hi-Z and Hi-Stencil. These early-rejects can be disabled both by some combinations of render states, as well as changing your depth test direction in the middle of the frame. When these early-rejects are not working your pixel shaders will execute for all pixels, even if depth or stencil tests kill the pixels. The result will be visually correct, but the official location for these depth and stencil tests is after the pixel shader.

Writing depth or stencil while testing depth or stencil will disable the early-reject for the draw calls doing this. This is sometimes unavoidable, but luckily only affects the specific draw calls that are setup this way.

On a lot of NVIDIA hardware, if you change the depth write direction (like I mentioned doing a pass of 'always' before 'lessequal' in order to the fix the Z-buffer on the XBOX), the zcull and scull will be disabled UNTIL THE NEXT DEPTH & STENCIL CLEAR. I expect this to be better or a non-problem with Geforce 280 series and newer, but haven't looked into it for sure. This also means you should always clear both at least at the start of the frame (and use the API's to do it, and not render a quad). This also makes the alternating depth lessequal/greaterequal every other frame trick to try and avoid depth clears a colossally bad idea.

The early-stencil test is very limited. On most hardware It pretty much caches the result of a group of stencil tests on some block size number of pixels and compresses it down to a a few bits. This means that using the stencil buffer for storing anything other than 0 and 'non-zero' pretty much worthless. And if you test for anything other than ==0 or !=0, the early stencil reject is not likely to work for you. It also means sharing the stencil buffer with a second mask is extremely difficult if you care about performance, and I definitely don't recommend trying it unless you can afford a second depth buffer with its own stencil buffer.
http://www.gearboxsoftware.com/
Lutz
Lutz
Wow, first of all a huge huge thank you for your answer. I really appreciate it. A lot of it sounds very familiar (like artists crying for more). I'm still digesting your post, but I'm going to have some questions. One is, what do you mean with "[color=#1C2837][size=2]render the gun with a viewport set to a range of 0/0.0000001 in order to get the gun to occlude the world". That range, is that the depth range? So you basically force depth=0?

One thing I can add: It is possible to read directly from the depth buffer in DX9 (NVidia 8xxx series and ATI 4xxx series I think). You have to create an INTZ texture ((D3DFORMAT)MAKEFOURCC('I','N','T','Z')), then get surface 0 from it (GetSurfaceLevel) and pass the surface into SetDepthStencilSurface before rendering. If you read from that texture, the value is going to be between 0 and 1 in homogeneous device coordinates. To get back view-space depth, you have to unproject it. This burns down to


// inverseRenderTargetSize = float2(1/renderTargetSizeX, 1/renderTargetSizeY)
float2 normalizedDeviceCoords = (VPos+ 0.5) * inverseRenderTargetSize;
float rawDepth = tex2D(DepthBufferTex, normalizedDeviceCoords).x;
float2 gamma = rawDepth * InverseProjectionMatrix._33_34 + InverseProjectionMatrix._43_44;
float viewDepth = gamma.x/gamma.y;

This is shader model 3 and up and VPos is the pixel position (as provided by SM3).

- Lutz
Zoner
Zoner

Wow, first of all a huge huge thank you for your answer. I really appreciate it. A lot of it sounds very familiar (like artists crying for more). I'm still digesting your post, but I'm going to have some questions. One is, what do you mean with "[color="#1C2837"]render the gun with a viewport set to a range of 0/0.0000001 in order to get the gun to occlude the world". That range, is that the depth range? So you basically force depth=0?


Yes. Unlike the XY dimensions of the viewport the projected depth of the geometry is normalized into the range specified by the viewport. I believe it needs to have a non-zero range so D3D or the hardware doesn't choke on division by zeros. This particular approach came from Epic when we updated our drop of Unreal. We had been using the stencil buffer approach previously.


One thing I can add: It is possible to read directly from the depth buffer in DX9 (NVidia 8xxx series and ATI 4xxx series I think). You have to create an INTZ texture ((D3DFORMAT)MAKEFOURCC('I','N','T','Z')), then get surface 0 from it (GetSurfaceLevel) and pass the surface into SetDepthStencilSurface before rendering. If you read from that texture, the value is going to be between 0 and 1 in homogeneous device coordinates. To get back view-space depth, you have to unproject it. This burns down to


// inverseRenderTargetSize = float2(1/renderTargetSizeX, 1/renderTargetSizeY)
float2 normalizedDeviceCoords = (VPos+ 0.5) * inverseRenderTargetSize;
float rawDepth = tex2D(DepthBufferTex, normalizedDeviceCoords).x;
float2 gamma = rawDepth * InverseProjectionMatrix._33_34 + InverseProjectionMatrix._43_44;
float viewDepth = gamma.x/gamma.y;

This is shader model 3 and up and VPos is the pixel position (as provided by SM3).

- Lutz


I'm not a fan of the vendor specific D3D hacks, though I have used them in the past (rudimentary access to NVIDIA register combiners in D3D8 for pre-shader 1.1 hardware comes to mind). It just makes compatibility testing a mess when you find out half the drivers out there are older than when the feature was added, or some subset of the vendor's cards do not provide the feature after the feature was created (ATI fetch4 comes to mind here), or a brand new latest gen piece of hardware no longer supports it. Depth buffer access is mostly a problem with D3D9 and older so should go away when we can finally drop it. Though that will be quite a while still, since Microsoft didn't get D3D 10level9 to support SM3.

Consoles can easily read depth textures, so it's certainly nice when you can use it, though a lot of times bundling the depth with other coherent data (with the normal in a deferred renderer comes to mind) can be much more efficient, as memory bandwidth repeatedly comes up as the bottleneck, as well as the quantity of of texture fetches.
http://www.gearboxsoftware.com/

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.