Skip to main content
GameDev.net gamedev.net
🔒 Locked

Optimization Help - XNA Hardware Instancing

Started by ChristOwnsMe Jun 20, 2012 at 10:01 AM 7 replies 1.7k views
Original Post
ChristOwnsMe
ChristOwnsMe
Hello. My program is rendering lots of cubes using hardware instancing and I am lagging even when I am not rendering that many cubes. I ran microsofts hardware instancing demo and I could render 50,000 cubes before I dropped to 30 fps,whereas mine I drop to 20 fps after 4,000.I am using chunked lod to divide the terrain into patches. Each patch has 2 vertex buffers and an index buffer. Does hardware instancing only work fast if I use less vertex buffers? Here is a screen. It is super zoomed in to the cubes so you cant see anything really. Any help or guidance on to what I could do to find the issue would be appreciated.


8519.png
Thanks.
Ashaman73
Ashaman73
You are talking about cubes, in the next sentence you switch over to terrain patches and finally you present an image where

you cant see anything really.

Could you provide more information. Have you extended the microsoft demo or have you written your own instancing code ? What happens if you turn off the terrain renderer ? Have you considered, that rendering terrain costs some time too ? How many patches are you handling ?`....
ChristOwnsMe
ChristOwnsMe
I have a terrain patch that renders cubes at points, instead of using vertices, but it still uses chunked lod to manage them. The picture was to just convey some text information about how much was being rendered. I implemented hardware instancing exactly like the demo, so I don't think that is the issue.I am rendering 37 patches and 3700 cubes when it begins to lag a lot. The only "terrain" is the cubes. I am making a minecraft lod planet, which is why I am using chunked LOD. And so instead of vertices in my patches, i have cubes at points in the patch.
Krypt0n
Krypt0n

Hello. My program is rendering lots of cubes using hardware instancing and I am lagging even when I am not rendering that many cubes.

do you mean input lag? that could identify that your GPU is fully loaded, so until your input is shown, some frames will pass, as the GPU command buffer might hold several frames. imagin you have 20fps with 5frames in the queue, additionally your game might have some processing time, it can accumulate to 300ms, this is where the lag comes from. it also identifies that instancing wouldn't really solve the problem, as your GPU is already fully loaded. Instancing is meant to overcome cpu limits to keep the GPU busy (which would be the case already).


I ran microsofts hardware instancing demo and I could render 50,000 cubes before I dropped to 30 fps,whereas mine I drop to 20 fps after 4,000.[/quote]
what do you do differently to the instancing demo? chances are high, you are not drawcall bound, but somewhere else. have you validated you're drawcall bound before you implemented instancing to optimize that? and if you did, how have you _exactly_ validated that this is/was the issue.


I am using chunked lod to divide the terrain into patches. Each patch has 2 vertex buffers and an index buffer[/quote]
terrain and instancing together? are you drawing the same terrain chunk several times?

Does hardware instancing only work fast if I use less vertex buffers? [/quote]
instancing is usually slower than drawing individual individual objekts, BUT it reduces the amount of draw calls a lot, so even if you pay with slightly slower drawing, you reduce the cpu cost.

if you are not limited by the drawcall count, you will end up with no speed up, maybe even a decrease in performance.
Nik02
Nik02
Instancing can be thought as a way of representing your meshes as first-form normalized database; you lose a bit of performance in the input assembler due to added fetch operation, but you save a lot of memory (and therefore total bandwidth) as you don't have to store or transfer redundant information.

Whether or not instancing is beneficial for you depends on the exact bottlenecks in your scenario.
Niko Suni
ChristOwnsMe
ChristOwnsMe
Yes, I am using "terrain" and instancing at the same time. I wan't to do a type of voxel LOD, and using normal chunked LOD was the only way I could think of to achieve this. Does anyone recommend any profilers to help me determine where my bottlenecks are? I have never profiled before,so finding where bottlenecks are is very new to me. Thanks for the responses.
Nik02
Nik02

  • PIX
  • GPU PerfStudio (AMD)
  • NSight (NVidia)

    I believe Intel also has GPU profiling tools, but I haven't used them.

    In addition, Visual Studio 2012 (Pro and up) support GPU profiling.
Niko Suni

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.