Skip to main content
GameDev.net gamedev.net
🔒 Locked

Deferred vs Forward shading

Started by trsh Feb 20, 2020 at 12:32 PM 13 replies 18.6k views
Original Post
trsh
trsh

Now I have been looking and reading into lots of examples, papers and could not come to conclusion which one is best to use. https://learnopengl.com/Advanced-Lighting/Deferred-Shading writes that 1847 lights would be never possible with forward rendering. Again there is this guy https://www.3dgep.com/volume-tiled-forward-shading/, who has a demo that runs thousands of lights with forward clustered + BHV. This guy again tells that both (Deferred & Forward) have advantages and disadvantages and mix them https://github.com/dtrebilco/lightindexed-deferredrender. So I'm kind of confused. I think for best performance I must go with clustering + BHV, but for the Deferred or Forward choice.. no idea. Maybe Both?

JoeJ
JoeJ

The problem of finding the lights affecting a pixel is the same for both, so any technique (BVH, binning lights to coarse tiles /cells, froxels etc.) is applicable for forward and deferred methods.
(Exception: Only with deferred you can draw the light volumes as polygons and accumulate the lighting, which was the first common practice to do deferred lighting, but seems much less common nowadays.)

But: with forward some GPU threads at the edges of triangles will do no work, while with deferred there are no such unused areas.
Conclusion would be that small triangles and costly lights gathering hurts forward more than deferred.

Downside of deferred is extra bandwidth to fill the GBuffers and restricted material options.

Both arguments make the best choice still depending on scene and materials, but not necessarily so much on the number of lights anymore.

trsh said:
Maybe Both?

For transparency support you may need a forward path in any case. So it makes sense to start with that, but keep a deferred path in mind.
In the end you might end up using deferred for most parts of the scene (e.g. static background with standard PBR material), and using forward for special materials (transparency, custom shaders, special effects…)

One more point to think about is support for raytracing. For this, you may need ability to gather affecting lights at any hitpoint, even offscreen.
From that perspective having a global acceleration structure for lights like BVH becomes more attractive, eventually.

Frantic PonE
Frantic PonE

One can note that there's a sexy third option called deferred texturing, that I don't believe has shipped yet in a project: https://therealmjp.github.io/posts/bindless-texturing-for-deferred-rendering-and-decals/

Basic idea is in the primary pass you read an indirection texture to the compressed material you want to use for rendering. A lot of the benefits of deferred, an absolute ton of triangles, and no restrictions on materials. Unfortunately you also (generally) need to output a bunch of extra odd parameters like UV gradients so it doesn't really end up saving bandwidth over deferred. Combined with the oddity of the whole process I'm not sure how popular it'll end up being, as the most obvious use case is with all virtualized textures as that just works right into deferred texturing pipeline, but only Idtech does that still.

JoeJ
JoeJ

Another option in between would be a thin gbuffer prepass with z, normal and roughness, then deferred lighting, then final full texturing and materials pass using the lit results in forward manner.

There are just too many options. Likely the best way to narrow it down is having some other technical restrictions and requirements.

trsh
trsh

JoeJ said:

Another option in between would be a thin gbuffer prepass with z, normal and roughness, then deferred lighting, then final full texturing and materials pass using the lit results in forward manner.

There are just too many options. Likely the best way to narrow it down is having some other technical restrictions and requirements.

And what would be performant for shadows?

MJP
MJP

If you're reading older articles about deferred rendering/shading, you have to keep in mind that deferred rendering came about before things like forward+ and clustered forward were introduced (mainly because those techniques rely on things like compute shaders and generalized buffer access in all shader stages, which were not at all a thing in the 2000's when deferred rendering started becoming popular). So when you see an older article like that make a comparison to forward rendering, they're really comparing with old-school forward rendering techniques. These generally fell into two categories:

  1. Single-pass rendering, with some fixed number of lights affecting each draw call (which might be dictated by hardware/API limits if using fixed-function lighting).
  2. Multi-pass rendering, where each draw is submitted multiple times with additive blending to sum the contribution from all light sources. In the worst case this means you do NumMeshes * NumLights draw calls.

Neither of these scaled up very well because they both tied your lighting granularity to your draw call granularity. If you wanted fewer draw calls with large meshes, you would either blow through your fixed lighting limit and/or need to draw that mesh many times to get the correct result. Deferred rendering decouples your geometry from lighting, which fixes that problem. Once more modern forward rendering techniques showed up they also solved that problem, just in a different way.

As for whether modern deferred vs. modern forward is better, I don't think there's a definitive answer on that one. Really nice-looking games continue to come out with both approaches, so I don't think you're going to be limited whichever route you choose. However JoeJ does bring up a great point in that you probably want a G-Buffer to do hybrid ray-tracing techniques, so if you're interested in that things then deferred may be a better choice.

trsh
trsh

@MJP Can you point me to some Modern forward and deferred lighting techniques / sources / papers?

JoeJ
JoeJ

Nice resources are ‘frame breakdowns’, e.g. https://aschrein.github.io/2019/08/11/metro_breakdown.html

http://www.adriancourreges.com/blog/2016/09/09/doom-2016-graphics-study/

http://www.adriancourreges.com/blog/2015/03/10/deus-ex-human-revolution-graphics-study/

… exist for many popular games.

trsh said:
And what would be performant for shadows?

The deepest rabbit hole of all?

I could link to many adventurous papers about efficient shadowed many lights, or approximate area light shadows, or compression of static shadows… if you have a certain interest.

But to me all this seems mostly unpractical / not worth it. (My biggest hope in RT is to get rid of shadow maps, and i don't think it's worth to invest in SM any longer. But not everybody agrees here.)

In practice, probably cascades for sun and some shadowed, some unshadowed local lights + shadow map cache makes sense mostly. Depends on the game, but i'm not really up to date what's the norm here actually.

Juliean
Juliean

The shading that I implemented is “Clustered Shading", based on this paper (http://www.humus.name/Articles/PracticalClusteredShading.pdf), from the guys that made Just Cause 3, etc…

The cool thing about clustered shading is that you can share your lighting-informating between forward and deferred rendering. The light lists you build are identical for both, making it extremely easy to have both forwrad and deferred passes. That is great for comparing the performance based on your scenario, as well as having forward-passes for ie. transparency if you decide to use deferred shading as the main render path.

trsh
trsh

@Juliean Do you have a public repo? And for shadows?

Juliean
Juliean

trsh said:
@Juliean Do you have a public repo? And for shadows?

No sorry, and it probably wouldn't help much as I implemented it in term of my engines API-abstraction layer which you'd need to know in order to understand the code :D
I also didn't get to shadows so far, concentrating on a 2D for now.

I mean, in case you are still interested, here is my code for processing the lights list on CPU (I didn't yet care so much for optimization of this step, so I didn't use any compute shaders for now): https://pastebin.com/FDqGvpGE

And here is the shader-code for interpreting this data (calculateClusteredLight is fed with data from eigther forward or deferred): https://pastebin.com/aJxzGv7M (the actual textures/cbuffer are as I said defined from inside the engine)

trsh
trsh

@Juliean Thanks. Will check out soon

D956
D956

I also implemented clustered shading based on the description in “idTech 666: The Devil is in the Details”.

The clustering is done on the CPU, SIMD optimised and multithreaded with one task per depth slice. The algorithm takes as input a list of items (lights, decals, probes, ...) each bounded by either an oriented bounding box or frustum.

I first compute the cluster bounds by projecting the item’s axis-aligned bounding box (AABB) and then cull by testing each cluster in the range against the planes of the item. This is done in normalised device coordinates (NDC) because there the clusters are just AABBs.

A cluster is culled when it lies completely in front of any one of the item’s planes. In order to test for this in NDC-space we first have to transform the planes:

For view-space plane \( \mathbf{f} (f_a, f_b, f_c, f_d) \), point \( \mathbf{p} (p_x, p_y, p_z, 1) \) and projection matrix \( P \):

\[ \mathbf{f}^\intercal \mathbf{p} > 0 \\ \mathbf{f}^\intercal P^{-1} P \mathbf{p} > 0 \\ (P^{-\intercal} \mathbf{f})^\intercal (P \mathbf{p}) > 0 \\ \mathbf{g}^\intercal \mathbf{p}’ > 0 \]

Since for all points in front of the view point \( p’_w > 0 \) we can divide by it and get:

\[ \mathbf{g}^\intercal (\mathbf{p}’ / p’_w) > 0 \]

Then the NDC-space cluster AABB with centre \( \mathbf{c} (c_x, c_y, c_z, 1) \) and half-extents \( \mathbf{r} (r_x, r_y, r_z, 0) \) (where \( r_x, r_y, r_z > 0 \) ) is in front of the plane if

\[ min[\mathbf{g}^\intercal(\mathbf{c}\pm\mathbf{r})] > 0 \]This expands to\[ (\mathbf{g}^\intercal \mathbf{c}) + min [ \pm g_a r_x \pm g_b r_y \pm g_c r_z ] > 0 \\ (\mathbf{g}^\intercal \mathbf{c}) - (\left| g_a \right| r_x + \left| g_b \right| r_y + \left| g_c \right| r_z) > 0 \]

Note that we can compute the left hand side once for each plane and then simply update it as we move from cluster to cluster:

f32 d0 = (g.a*c.x + g.b*c.y + g.c*c.z + g.d) - (abs(g.x)*r.x + abs(g.y)*r.y + abs(g.z)*r.z);
f32 dx = 2*g.a*r.x;
f32 dy = 2*g.b*r.y;

for(u32 y = y0; y <= y1; y++)
{
  f32 d1 = d0;
  d0 += dy;
  for(u32 x = x0; x <= x1; x++)
  {
    if(d1 > 0)
      cluster in front of plane
    d1 += dx;
  }
}

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.