Skip to main content
GameDev.net gamedev.net
🔒 Locked

Minimizing draw calls improves performance?

Started by programci_84 Jan 29, 2011 at 10:39 PM 5 replies 5.6k views
Original Post
programci_84
programci_84
Hi all,

I've heard that one of the tricks of improving performance is minimizing draw calls.

First question: Is that really true?

If so,
Second qustion comes:
Suppose I have a mesh consists of approx. 600,000 tris and contains 25 subsets. 10 of 'em are using, say, Material1; 8 of'em are using, say, Material2 and the rest are using Material3. Now I have 2 options:
* For each material, rendering subsets which are using that material (Is this called "Sorting by Material" ?) .
* Building one subset from subsets which are using the same materials (After that, I'll get 3 "big" subsets instead of 25, and number of draw calls will be decreased, correct?) .

Which option will be better?
Also, is it possible to perform 2nd option every frame (because some material info can be changed real-time) ? If so, how? By locking and unlocking vertex/index buffers, or something like that? Details, please.

Thx. in advance.
-R
There's no "hard", and "the impossible" takes just a little time.
Erik Rufelt
Erik Rufelt
It can be true. In your case you have nothing to gain as 25 draw-calls is nothing for 600k tris. If your model had only 600 tris however, and you draw 1000 copies of it, then you should look into options for minimizing the calls.
smasherprog
smasherprog
Question 1:

Yes, the reason is because each call to the draw function has a fixed cpu cost associated with it. What this means is that the cpu has to do a little bit of work to complete the function call. In fact, on modern graphics cards, if you do not submit at least 400 polys per draw call, you are spinning your tires. This does not count for instanced calls, in an instanced call you would want the total number of polys to be somewhere above that number. Now, this number is not exact, and it will vary +-100. If you submit a draw call with 100 polys, tne video card will finish the work before your cpu can submit another draw call. This is called starving the GPU. This is why you should always try to decrease the number of draw calls.

Question 2:

For each different material, you will have to issue a separate draw call regardless. Unless you can fit the textures into a texture array and pass along the texture index' in a separate vertex stream. But, like erik said, 600k triangles and 25 draw calls is fine. if you can decrease the draw calls, do it. But, dont put in a bunch of work to get a 1% speed increase, its not worth it. Use your time efficiently.
Wisdom is knowing when to shut up, so try it.
--Game Development http://nolimitsdesigns.com: Reliable UDP library, Threading library, Math Library, UI Library. Take a look, its all free.
Zoner
Zoner
The bulk of the cost of a draw call is processing the preceding state changes. If you are just trying to draw different subsets of the same mesh its pretty fast, though if you abuse this you will still feel it.

That said a good plan for being 'nice to the hardware' is to call as few d3d functions as possible.
http://www.gearboxsoftware.com/
programci_84
programci_84
Thank you all, for your fast replies.

I'm also wondering is it a good idea to use a big vertex buffer for combining subsets which are using the same material, on every frame ?

For example;
void MyRenderFunc()
{
//...

fatVertexBuffer1->Lock( ... );
for each subset using Material1
Add these subsets' vertices to fatVertexBuffer1;
next
fatVertexBuffer1->Unlock( );

fatVertexBuffer2->Lock( ... );
for each subset using Material2
Add these subsets' vertices to fatVertexBuffer2;
next
fatVertexBuffer2->Unlock( );

fatVertexBuffer3->Lock( ... );
for each subset using Material3
Add these subsets' vertices to fatVertexBuffer3;
next
fatVertexBuffer3->Unlock( );

DrawPrimitive (fatVertexBuffer1);
DrawPrimitive (fatVertexBuffer2);
DrawPrimitive (fatVertexBuffer3);
}


Is this possible? Does it save our time, or not?
Thanks.
There's no "hard", and "the impossible" takes just a little time.
Zoner
Zoner
I would say that the next biggest 'sin' in an engine is moving memory unnecessarily.

If that geometry is static it is either better off being kept in separate buffers, or merged once and left alone. Switching vertex and index buffers are also relatively cheap as long as the vertex decl (d3d9) or input layouts are left alone (d3d10/11), and so are textures, though once you make too many exceptions it starts looking like un-optimized rendering code again
http://www.gearboxsoftware.com/
Jason Z
Jason Z
The CPU cost can also be reduced by utilizing multithreaded rendering in D3D11. By submitting the state changes, resource manipulations, and draw calls on deferred contexts that run in separate threads on a multicore CPU, you can effectively reduce the average cost of submitting the work to the GPU. I don't know if you are using D3D11 or not, but it can have a dramatic improvement in performance if you have the right conditions in your scene. In recent testing, I have seen from 46-68% frame time improvement depending on the number of objects in the scene...

With that said, reducing draw calls only helps you if you are CPU bound. For example, you could submit one hundred triangles per draw call for 10000 triangles total, and if you do sufficiently complex calculations on the triangles then reducing the number of draw calls won't improve your speed at all. Each situation and scene is different, and you should strive to make your engine flexible enough to be able to switch strategies based on what you are rendering.
Jason Zink :: DirectX MVP   Direct3D 11 engine on CodePlex: Hieroglyph 3 Direct3D Books: 

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.