Skip to main content
GameDev.net gamedev.net
🔒 Locked

Occlusion Culling

Started by Dirge Apr 27, 2003 at 1:48 PM 35 replies 29.5k views
Original Post
Dirge
Dirge
This topic just cropped up in the Yann worship thread but I thought I''d start a new thread to discuss it. Yann mentions that with Occlusion Culling he is able to get a general and reliable algorithm that applies for both indoor and outdoor culling. What I''m interested in is what techniques he, and others here are using. Also how do most people merge this technique with an Octree or other spatial data structures (unlike Hierarchical Z-Buffering which relies on it). In other words point or cell based? Image, object or Ray Space? Hardware or Software based? Occlusion Horizens, or Shaft Occlusion Culling? Hierarchical Z-Buffering or HOM? etc... What have you guys found easiest to implement and practical, yet effective? Also, you''d have to sort front-to-back to have your occluders work at all, how do you deal with semi-transparent objects then? What other issues might come up because of this? "Love all, trust a few. Do wrong to none." - Shakespeare www.CodeFortress.com
"Artificial Intelligence: the art of making computers that behave like the ones in movies."www.CodeFortress.com
duhroach
duhroach
Yann uses a modified HOM algorithm for his Culling system. That along with his ABT yields a pretty nice system.

The main prospect to get OC working with your spacial subdivision algorithm, is that each of your nodes contain a chunk of geometry unique to that node. Therefore, any nodes visible in your frustum can easily be depth sorted as an entire node, and sub sorted if you desire. From there, walking node by node, rendering Cullers etc becomes quite easy.

I actually use a modified version of the cPLP algorithm (Check out Game Programming Gems 2/3) Additionally, for models that aren''t 100% culled or unculled, models are broken down into sub models via an additional ABT. Which creates a tightly bounded box around given chunks of the model. This allows me to do sub chunk testing, so i only have to render what''s REALLY visible.

~Main

==
Colt "MainRoach" McAnlis
Programmer
www.badheat.com/sinewave
Dirge
Dirge
Anyone else?
"Artificial Intelligence: the art of making computers that behave like the ones in movies."www.CodeFortress.com
v71
v71
software depth map here

Yann L
Yann L
I can't speak for others, but I have made highly satisfying experiences with an essentially very simple approach. Perhaps I should quickly outline how I ended up with that method, and why I think it is the best choice on current hardware.

I started, as surely many other people, with a portal system. Unfortunately, the inherent complexity of our levels made the portal approach very difficult. Automatic portalization was no option, so the artists needed to place the portals manually. The portals could have any coplanar shape. We used a portal/anti-portal approach, connecting sectors in both directions. A special sector, the NPS (non-portalized space), grouped all the geometry that did not fit into a specific sector. At runtime, the system would recursively go through the list of visible portals, and clip them to yield depth dependent visibility masks for each sector. So far, so good. Performance was OK, culling was good too.

But our artists were everything else than happy. First, placing all the portals manually is a pain in the ass. And second, being the most important point, it limited them tremendeously in their creativity: everything they modelled, had to fit into this rigid portal/sector concept. Everything that didn't was thrown into the NPS. Now, on our first scene, we (the programmers) had quite a shock, when we compiled the level: the artists had put around 90% of the level geometry into the NPS ! It was just not portalizeable. If you have a large forest, how do you portalize it ? Or a village, seen from the outside, where do you put the portals ?

We needed an alternative. We first tried a kind of full 3D version of a PVS: precomputed volumetric visibility lists, using fuzzy projections. That works very well, but takes an insane amount of memory for large scenes. It also takes too much preprocessing time.

So I tried shaft culling. Basically, very easy, and fully object space. But the time used to cull objects against thousands of shafts was far too much. Furthermore, since occluder fusion is very hard to achieve with this technique, culling results were far below our expectations. And not to forget: being a geometrical algorithm, it is prone to mathematical instabilities due to floating point accuracy problems. Not good. And the idea to merge the shafts into a fusioned occlusion volume in realtime, resulted in a total nightmare.

And so I ended up using HOMs: image space method, easy to generate, totally stable under all circumstances, and implicit occluder fusion. The first tests delivered excellent results, both in terms of occlusion map generation, as well as culling efficiency. I modified the well known Zhang-algorithm a little, exclusively using a high accuracy 32bit depthmap (ignoring the coverage map). The approach fitted perfectly in our spatial ABT structure: since the ABT offers tightly bound objects in AABB's, and feature local vertex/index pools, they can directly be culled against the HOM, by a simple 4x4 transform/project. The culled nodes are simply not send to the render scheduler. Extremely efficient.

Now, a few months ago, I decided to reimplement the occlusion system on the GPU, using this HP/NV occlusion query stuff. Well, to make it short, I trashed the code pretty fast. Due to the bubbles introduced into the command stream (even by using multiple deferred queries with the NV extension in OGL), it was slower than the software approach. Esp, if the later uses CPU/GPU concurrency.

So our current system renders the next frame's occlusion map on the CPU, in parallel to the GPU rendering the currently scheduled frame. Basically, I get the occlusion map for free. Rendering is done by using a very fast software rasterizer, written in hand optimized ASM. Same for the query system (comparing a projected AABB against the HOM). I'm very happy with that configuration: activating both rendering and querying will generally drop the framerate by about 10%, but considering the fact that the visible geometry is often reduced by amount as high as 95%, it's more than worth it.


[edited by - Yann L on April 30, 2003 11:41:37 AM]
SantaClaws
SantaClaws
Yann:

Are you saying you don''t use the method of opacity/transparency threshold based testing against HOMs that Zhang generates, and only use a higher resolution depth estimation buffer, or something else?
Zemedelec
Zemedelec
quote:
Original post by Yann L


And how the HOMs are being generated? Having one (ABT) node, we know its BV, i.e. test model. But what is its write model (occlusion model)?
It isn''t the original geometry resterized by CPU, I am pretty sure...
blue_knight
blue_knight
I think he uses virtual occluders. Simplified geometry that works well as an occluder.
Yann L
Yann L
quote:

Are you saying you don''t use the method of opacity/transparency threshold based testing against HOMs that Zhang generates, and only use a higher resolution depth estimation buffer, or something else?


That''s right, I only use a zbuffer for the test. Since I''m rendering the occ-map in software, rasterizing only a single 32bit channel is almost twice as fast as when rendering the additional opacity channel. I can still take advantage of non-conservative culling, by using the ratio of failed to succeeded depth tests when comparing the projected AABB against the HOM.

quote:

And how the HOMs are being generated? Having one (ABT) node, we know its BV, i.e. test model. But what is its write model (occlusion model)?
It isn''t the original geometry resterized by CPU, I am pretty sure...


Depends. As blue_knight mentioned, I use virtual occluders (in object space) for the general occlusion (for static geometry, as well as for dynamic). You can see it as a kind of occlusion skin. For problematic things like trees, vegetation, and all complex ''fuzzy'' objects that do not have a large occlusion footprint by themselves (but generate a considerable fusioned occlusion, think of a forest, for example), I use a low-poly LOD version of the object (with conservative volume). For a tree, that would be a set of alpha patches. Now, this can sometimes lead to false positives (eg. a tree can sometimes dissappear, if lots of small foliage from another tree is infront of it), but in practice you don''t really notice that (it''s only a few pixels). And it only happens with complex organic objects, that do not have an adequate virtual occluder.
treething
treething
Do the artists create the occlusion geometry by hand, or is it created automatically?

I''ve been thinking of writing a similar system for years (ever since i read about dPVS/Umbra), but never get the time for it.

Do you use geometric tests to supplement the HOM, as in dPVS? I vaguely remember them using ray-casts to track visible points on an object (corners of the bounding box), as a positive early out for the visibility test.
jmchambs
jmchambs
So essentially you''re leaving the H off HOM and just making one occlussion map, or depth map or whatever. Have you tried generating a hierarchy to see if you get any speed up in occulsion testing? Or are filtered opacity maps not useful in your system?

Which indirectly leads me to my real question. How do you determine the occulders? I read that chapter of Zhang''s thesis and it left me feeling like there has to be a better way. But of course as Zhang mentions "The optimal occluder set is exactly the visisble portion of the model". Do you just use distance based selection? Do you use the potential for temporal cohence at all? (that is one area where it seems improvements could be made)

One other thing. Are the low LOD models used for occluder skins the same as the models used in low detail rendering? i.e. Do you precompute models specifically for occullsion?
CoffeeMug
CoffeeMug
I believe occlusion skins have to be precomputed separately from low detail models used for rendering. Occlusion skins have to have conservative volume (meaning they can never be larger then the original mesh), while a low LOD model could probably benefit from being free from this requirement.
jmchambs
jmchambs
But would the benifit of not being volume restricted, out weigh the consequence having an additional LOD (or multiple LOD''s) specifically for occlusion?
Heretic
Heretic
Yann: I have to admit - you have convinced me to use occlusion culling However first Im going to use portals and then add OC and make it hybrid. I think Ive got a nice portalization method that can come almost for free and help the hsr.
Ive been thinking however about something else... what do ya think about C-buffers ? How do they perform ? Isnt it actually faster to use them, in some form they can store pixels as single bits thus reducing memory overhead... They could be hierarchical as well as the HOM and I`ve been wondering about which one should I use... Help me ?
Yann L
Yann L
quote:

Do the artists create the occlusion geometry by hand, or is it created automatically?


Automatic, most of the time. On some models, however, the automatic system can fail. That''s where the artists have to jump in.

quote:

Do you use geometric tests to supplement the HOM, as in dPVS? I vaguely remember them using ray-casts to track visible points on an object (corners of the bounding box), as a positive early out for the visibility test.


No, I exclusively do image space tests. But since it''s a hierarchical buffer, the early rejection case is very fast (operates on low resolution z-levels).

quote:

So essentially you''re leaving the H off HOM and just making one occlussion map, or depth map or whatever.


No no, don''t get me wrong: I''m still using a hierarchy. But without the coverage value, only with z. The level 0 z-map is just a standard 32bit depth buffer. At higher levels, each pixel represents a z range, the minimum and maximum z value covered by the child pixels.

quote:

I believe occlusion skins have to be precomputed separately from low detail models used for rendering. Occlusion skins have to have conservative volume (meaning they can never be larger then the original mesh), while a low LOD model could probably benefit from being free from this requirement.


That''s correct, the occlusion skins are separate.

quote:

But would the benifit of not being volume restricted, out weigh the consequence having an additional LOD (or multiple LOD''s) specifically for occlusion?


Occlusion skins typically have very few faces, as they are being software rendered. The additionally required memory resources are minimal.

quote:

Yann: I have to admit - you have convinced me to use occlusion culling


Good

quote:

Ive been thinking however about something else... what do ya think about C-buffers ? How do they perform ?


You mean for the software occlusion map rendering ? I don''t know, I never tried them. I have a doubt that they would be very effective (the occlusion rendering pass is far from being a bottleneck), but you could possibly get a speedup if you have lots of occlusion geometry with high depth complexity.

quote:

They could be hierarchical as well as the HOM and Ive been wondering about which one should I use... Help me ?


Start with standard HOMs, and once that works, profile your code. If your rendering/testing phase really turns out to be a performance problem, then you can still implement them. Otherwise, you''ll just waste your time.

quote:

1. For Yann mainly... Using a software simplified Zhang HOM, where would you estimate that using GPU occlusion queries would out perform CPU based queries? For example, approximate processor speed at which using the CPU would be slower. Just your best guess.


Woah, very hard to say. It''s more than a simple function of the CPU speed, as FIFO bubbles do not always react as you expect them to. It''s more a synchronisation issue between the GPU and CPU. And it obviously heavily depends on the performance of your SW rasterizer. What is the target hardware of your product ? I tend to see it like this: if the CPU is too slow to handle occlusion culling, then it will be too slow to handle the game. I have tested my engine on systems as low as an Athlon 600Mhz. And even on this config (+GF4 Ti4400, iirc), the CPU driven culling was still so fast, that it wasn''t worth introducing the bubble that comes with HW culling.

quote:

2. When using DX9/XBox deferred occlusion queries, you can render one frame, wrapped in a query, and then next frame omit a mesh if it''s occlusion results indicate 0 for few pixels drawn. What sort of logic would you use to re-include a mesh into the scene? You''d need to perform a render to do occlusion tests, so the best this method seems to offer is to skip complex pixel effects and multiple passes until the mesh becomes visible. Am I missing something?


Yeah, that''s the classical GF3 style deferred OC test. With that method you can actually avoid bubbles (or better: push them to the end of the frame, where they aren''t that bad). But keep in mind, that this occlusion result was only valid for the last (already rendered) frame. As soon as the camera or the environment moves, the query results from the last frame might not be valid anymore. You could play with time coherency, but that''s risky: you might quickly get ugly popping artifacts. So basically, it''s mostly a trick to get the framerate up, when the camera and general environment are static. People do everything to get those 60fps minimum, needed to get MS approval for your Xbox product.

You can of course render the occlusion for the next frame, after drawing the current one, and use the query results to cut expensive perpixel stuff on the next frame. That''s entirely possible. But this again will introduce a bubble.

quote:

3. Regarding DX9 occlusion queries, nVidia claims some current card, and most future cards will support it. I can''t find any real info on which cards support what. Can anyone comment on which cards have it? Could you check your caps, and post your card type and whether it supports it? Is this GeForceFX, or Radeon 9700 and up, or do some earlier cards support it?


No idea about DX9, but on OpenGL, deferred hardware occlusion queries are supported on GF3 upwards, and on Radeon 7500+.
jmchambs
jmchambs
Still wondering what methods people use to choose occluders, since most OC schemes require a good method to get good results. What type of precomputed information do people use and how do you use it for dynamic occulder determination.

I''d imagine you can use a good geometry tree to determine whether or not individual leaves may make good occluders.

Do people use time coherence in estimate the occluder set for the next frame?
Heretic
Heretic
Yann: So whats the difference between your method and the hierarchical z-buffer ?
RandomLogic
RandomLogic
Hi,

When talking about a SW rasterizer. Does this include just what seems obvious to me in order to create the mask, i.e. transformation, line drawing and poly filling, or is there more to it?

Also just out of curiousity what methods (known algo''s???) are used for the line/fill tasks?

And one last... I assume the map is rendered to a 2D array in the main memory.. true or false?
There aren't any stupid questions,Only stupid people.
CoffeeMug
CoffeeMug
Actually it''s simpler then that. Since you''re only interetsed in the depth buffer the software rasterizer only handles drawing to a depth buffer. You don''t have to worry about coloring.
Zemedelec
Zemedelec
CoffeeMug:
Well, nobody mentioned colouring/texturing - it is obvious one do ot need ''em, isn''t it?
And if there is a way to use c-buffer over z-buffer all things get very fast...

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.