Skip to main content
GameDev.net gamedev.net
🔒 Locked

Only 12 Enemies, And My Fps Drops To 30, Why Is That?

Started by Heelp Jul 31, 2016 at 10:44 PM 31 replies 8k views
Original Post
Heelp
Heelp
Guys, I have 7 animated enemies. And my fps is 62 for now( not capped). But when I add 5 more enemies and make them 12, my fps drops to 31. I traced the problem and I finally found it, it's my BoneTransform() function, which fills my vector of TransformMatrices that I use in the vertex shader in order to animate the skeleton. But it rapes my CPU. ( when I comment the BoneTransform() function, framerate goes from 30 to 166!( sometimes jumps between 166 and 200 ). And I kind of stole most of the function from a tutorial on skeletal animation, and I'm sure it's pretty optimized, so there must be some other reason.

[attachment=32770:lowfps.gif]

I used some models from World of Warcraft. And the interesting thing is that I have the game, and when I play it( when I play WoW ), I can have 20 players around me, and my fps is great, but when I add the same models in my own game, my fps drops like crazy and it's 10 times slower than the original game, why? ( bear in mind that I haven't even loaded any map, I just spawn 12 enemies walking on air, and my cpu runs like a fat truckdriver, wtf is that?? ).
Servant of the Lord
Servant of the Lord

Enemies shouldn't "have" a deltaTime, you should just pass them the current deltaTime to their update function.

(right now it looks like you are mixing your updating, player input event processing, AI thinking, and rendering, all in one function)

I used some models from World of Warcraft. And the interesting thing is that I have the game, and when I play it( when I play WoW ), I can have 20 players around me, and my fps is great, but when I add the same models in my own game, my fps drops like crazy and it's 10 times slower than the original game, why?

Because it's not about what data you're loading, it's about how your code uses it. Well, okay, it's about the data and the code working together.

Your code and WoW's code is different, and thus your framerate and WoW's framerate is different.

(Make sure you don't use WoW's models in any copy of your game you distribute publically, btw - that's copyright infringement)

Heelp
Heelp

Thanks for the answer.

But!! :)

The way I see it, there are two ways.

First way: Pass the deltaTime to each function in the Enemy class. The problem with this method is that if I have 100 movement functions, I need to pass deltaTime 100 times every frame.

Second way: Pass the deltaTime in the class as a variable each frame. The pros are that I pass the deltaTime only once per frame and I can use it by 100 functions if I want to. The second way sounds better, because you pass the variable only once and use it as much as you want.

And about the fps drop, what would you suggest? I mean, what should I do, I'm pretty sure I have loaded the skeletal animation correctly, because most of the code is stolen from different tutorials, and I don't know what's wrong.

I think the problem is that when I start the game, it uses only 1 core on my laptop, but when I start WoW, all 4 cores are used, can this be the problem?

frob
frob



I have 7 animated enemies. And my fps is 62 for now( not capped). But when I add 5 more enemies and make them 12, my fps drops to 31. I traced the problem and I finally found it, it's my BoneTransform() function, which fills my vector of TransformMatrices that I use in the vertex shader in order to animate the skeleton. But it rapes my CPU.

First step, use your profiler and make sure you are calling it the right number of times. A common error is to call functions more often than needed.

Second step is to get all the performance numbers for the calls that you can using your profiler. Frames per second is a nearly useless number for profiling, you need specific times of specific functions in nanosecond or microsecond resolution.

If it is still running that slow and you don't know how to improve any specific performance number, share the performance problem, your profiling numbers and counts, and the source code in the appropriate area of the site (Graphics Programming, OpenGL/Vulkan, or Direct3D) that matches your code. If you've got more than 50 or so lines of code, consider using a paste site rather than doing a source dump in the discussion forums.

Heelp
Heelp

Ok, thanks frob, I will definitely search for a profiling tool because I haven't used one so far. Seems to me that this will take time to fix, so I will leave it for some later moment.

Another question came to mind and I don't want to make another post.

I loaded a very simple map for my paintball game. The map has a floor with a couple of walls and that's it. The floor is perfectly flat and the walls... well, I copied the floor, rotated it to 90 degrees and I made the walls with it.

And I need a very, very simple collision detection for this map. The only two options that come to mind are:

1.Make AABBs for the floor and the walls and do checks every frame.

2.Take a picture of the scene from above and store the depth values in a framebuffer and somehow use them to decide if the player is going to collide.

Is there something else I can do, and if not, what option should I choose from these two?

Kylotan
Kylotan
You should start a new topic to ask a different question.
Scouting Ninja
Scouting Ninja

It reads like you hit your geom limit, some times known as a object limit.

Lucky it is easy to test, double the poly count of the animated model and note the frame rate, then use a very low(100 polygon) animation model an note the frame rate.

If you still run at the same frame rate with low polygon models and high polygon models, give or take a frame or two, then it's the geom limit.

If you get low frame rate no matter the poly count it is often the geom limit or the shaders in my experience.

PC graphic cards can only render so much objects at real time, this is known as the geom limit or object limit. You can batch models into one model to use less objects or you can use instancing.

For animation objects you want instancing as dynamic batching can be very hard and unpredictable.

The problem is that this can be many things from draw calls to bad programming, modeling and many other things.

ferrous
ferrous



First way: Pass the deltaTime to each function in the Enemy class. The problem with this method is that if I have 100 movement functions, I need to pass deltaTime 100 times every frame.

Passing a single float, even 100 times, is very cheap, if not free. You're pre-optimizing and obfuscating code without any real gain.

Heelp
Heelp

ferrous, true story man, I always forget to think before optimizing..

Scouting Ninja, I don't think I've hit any limits because in the original game my pc can handle 10 times more enemies and everything is ok, but I will give it a try, I just need to change the models poly count with blender, thanks for the idea by the way.

Heelp
Heelp

Guys, I solved it.


const aiNodeAnim* Model::FindNodeAnim( const aiAnimation* pAnimation, const string NodeName )
{
    for ( uint i = 0 ; i < pAnimation->mNumChannels ; i ++ )
    {
        const aiNodeAnim* pNodeAnim = pAnimation->mChannels[i];

        if ( string( pNodeAnim->mNodeName.data ) == NodeName )
        {
            return pNodeAnim;
        }
    }

    return NULL;
}

This function here swallows 250 of my fps per second( from 380 to 30 ) and this is the function that finds the proper animation for every node in the model. Basically it counts from 0 to 120( in my case ), and for every loop it does a string comparison in order to find the proper animation for the node. Can you believe it? I still can't. This function is placed in a recursive function called ReadNodeHierarchy() that reads all the nodes matrices and calculates the interpolation and consequently, the final transformation matrix.

Here is the function:


void Model::ReadNodeHeirarchy( float AnimationTime, const aiNode* pNode, const Matrix4f& ParentTransform, int currentAnim )
{
    string NodeName( pNode->mName.data );

    const aiAnimation* pAnimation = this->scene->mAnimations[currentAnim];

    Matrix4f NodeTransformation( pNode->mTransformation );

    const aiNodeAnim* pNodeAnim = FindNodeAnim( pAnimation, NodeName ); //Only this function swallows 250 fps, believe it or not.

    if ( pNodeAnim )
    {
        // Interpolate scaling and generate scaling transformation matrix
        aiVector3D Scaling;
        calcInterpolatedScaling( Scaling, AnimationTime, pNodeAnim );
        Matrix4f ScalingM;
        ScalingM.InitScaleTransform( Scaling.x, Scaling.y, Scaling.z );

        // Interpolate rotation and generate rotation transformation matrix
        aiQuaternion RotationQ;
        calcInterpolatedRotation( RotationQ, AnimationTime, pNodeAnim );
        Matrix4f RotationM = Matrix4f( RotationQ.GetMatrix( ) );

        // Interpolate translation and generate translation transformation matrix
        aiVector3D Translation;
        calcInterpolatedPosition( Translation, AnimationTime, pNodeAnim );
        Matrix4f TranslationM;
        TranslationM.InitTranslationTransform( Translation.x, Translation.y, Translation.z );

        // Combine the above transformations
        NodeTransformation = TranslationM * RotationM * ScalingM;
    }

    Matrix4f GlobalTransformation = ParentTransform * NodeTransformation;

    if ( boneMapping.find( NodeName ) != boneMapping.end( ) )
    {
        uint BoneIndex = boneMapping[NodeName];
        boneInformation[BoneIndex].FinalTransformation = GlobalInverseTransform * GlobalTransformation *
                                                    boneInformation[BoneIndex].BoneOffset;
    }

    for ( uint i = 0; i < pNode->mNumChildren; i ++ )
    {
        ReadNodeHeirarchy( AnimationTime, pNode->mChildren[i], GlobalTransformation, currentAnim );
    }
}

This magical statement gulps all my CPU power: const aiNodeAnim* pNodeAnim = FindNodeAnim( pAnimation, NodeName );

This is because the readNodeHierarchy() function is recursive. For example you have one transformation matrix for the fingers, but then you need to multiply by the arm transformationMatrix because the arm moves the fingers too, and then the body moves the arm which moves the fingers and so on and so on. And every time that happens, the findNodeAnim() function counts from 0 to 120( mNumChannels) in order to find the proper animation based on the node's name and it does string comparison and some other crazy stuff million times per second.

And this was the code from this tutorial: http://ogldev.atspace.co.uk/www/tutorial38/tutorial38.html

It does the job for a tutorial, it is readable, but it is very unoptimized. What I did is to cache all the animChannels' indices with their proper bone in a map container, now the same 12 enemies run on 100 fps instead of 30 fps, so 70 fps gained by caching all the animations.

here is the fps with 24 enemies.

[attachment=32776:gamefixed20fps.gif]

I wonder what crazy fps gain can be made if I cache the interpolation matrices too?

jbadams
jbadams
I edited the "solved" out of your topic title; we generally discourage doing that, as it makes it less likely people will continue to read and respond to your topic, which might otherwise still have valuable conversation - you even ended with a question that may not be answered if noone checks.


Hopefully code taken from tutorials will be correct (sadly not always the case), but it's very commonly not optimised, and when glueing together code from different tutorials you might also introduce inefficiencies yourself: never assume your code will be optimised because it's from a tutorial, and always take actual measurements with a profiler rather than guessing blindly. :)
- Jason Astle-Adams
Heelp
Heelp

jbadams, sorry, I didn't know that, but it sounds logical, I agree.

Frames per second is a nearly useless number for profiling, you need specific times of specific functions in nanosecond or microsecond resolution.

frob, I'm still wondering, why fps would be useless number for profiling?

Alberth
Alberth

fps measures the entire loop, which means you only know total execution time.

Profiling gives information how much time each part (each function, or even each line) takes. Obviously, the total should be the same as the fps time, but in addition, you get a direct pointer to what code is causing the problem, rather than "somewhere in the loop I use too much time".

Lactose
Lactose

Also, "losing/dropping" x frames per second is a relative thing.

If you're at 100 FPS normally, and drop 20 FPS, your frametime goes from 10ms to 12.5ms (2.5ms increase).

However, if you're at 40 FPS normally and drop 20 FPS, your frametime goes from 25ms to 50ms (25ms increase).

While for the end product the frame rate is probably a more important metric, while developing and finding bottlenecks it's mostly useless.

Hello to all my stalkers.
Heelp
Heelp

Guys, very big thanks to all for all the helpful answers, this morning I couldn't load 12 enemies, now I'm loading 70, I'm very happy.

Lactose, I haven't thought about that, and you are right, basically the bigger the fps is, the smaller the difference in milliseconds is.

I was just wondering why my fps would jump from 100 to 111 every time, now when I think about it, 1000ms/100 = 10ms and 1000ms/111 = 9.09ms, and only the rounding error miscalculates the fps with 11!

EDIT: Alberth, I just installed some free profiler called VerySleepy, and it told me that printf takes 99% from the thread, and I checked all my code, I didn't have any printf functions at all.

But what I did have is "snprintf" and I guess verySleepy counts snprintf() at least as heavy as printf(). Luckily, I need this function to be executed only in the beginning of my game, not every frame, so I moved it out of the loop and fps jumped from 100 to 166! This snprintf is a very heavy function....


void MMORPG::getBoneLocation(Shader& shader)
{
    for (unsigned int i = 0 ; i < 100 ; i++)
    {
        char Name[128];
        memset(Name, 0, sizeof(Name));
        snprintf(Name, sizeof(Name), "gBones[%d]", i);
        boneLocation[i] = glGetUniformLocation( shader.programID, Name );
    }
} 

EDIT: Guys, VerySleepy continues to show that printf takes 90-100% of some thread. I checked all parts of my source code, I even tried commenting everything, fps went crazy and the profiler still says that printf works in background somehow, although I checked multiple times that I don't have anything that says 'print' in my whole project.

First I thought that the reason is that my project is set up as a console application in codeblocks, but I changed it to a GUI application and it still cries about printf, I profiled the printf, and the only thing it shows is where the function is declared in stdio.h. Any ideas?

SimonForsman
SimonForsman

Guys, very big thanks to all for all the helpful answers, this morning I couldn't load 12 enemies, now I'm loading 70, I'm very happy.

Lactose, I haven't thought about that, and you are right, basically the bigger the fps is, the smaller the difference in milliseconds is.

I was just wondering why my fps would jump from 100 to 111 every time, now when I think about it, 1000ms/100 = 10ms and 1000ms/111 = 9.09ms, and only the rounding error miscalculates the fps with 11!

EDIT: Alberth, I just installed some free profiler called VerySleepy, and it told me that printf takes 99% from the thread, and I checked all my code, I didn't have any printf functions at all.

But what I did have is "snprintf" and I guess verySleepy counts snprintf() at least as heavy as printf(). Luckily, I need this function to be executed only in the beginning of my game, not every frame, so I moved it out of the loop and fps jumped from 100 to 166! This snprintf is a very heavy function....


void MMORPG::getBoneLocation(Shader& shader)
{
    for (unsigned int i = 0 ; i < 100 ; i++)
    {
        char Name[128];
        memset(Name, 0, sizeof(Name));
        snprintf(Name, sizeof(Name), "gBones[%d]", i);
        boneLocation[i] = glGetUniformLocation( shader.programID, Name );
    }
} 

EDIT: Guys, VerySleepy continues to show that printf takes 90-100% of some thread. I checked all parts of my source code, I even tried commenting everything, fps went crazy and the profiler still says that printf works in background somehow, although I checked multiple times that I don't have anything that says 'print' in my whole project.

First I thought that the reason is that my project is set up as a console application in codeblocks, but I changed it to a GUI application and it still cries about printf, I profiled the printf, and the only thing it shows is where the function is declared in stdio.h. Any ideas?

it could be any function that outputs data to stdout (the console / terminal / IDE), a lot of such functions could be implemented as wrappers around printf), normally stdout will be line buffered when it is tied to a terminal (each newline will cause the buffer to be flushed which is a fairly slow operation), if you need to push large amount of text to the console you should flush manually when appropriate (once per frame, between levels, when the app closes, once the buffer is full, or whatever suits your application)

if you use c++ you can call std::ios_base::sync_with_stdio(false); and then avoid using std::endl (just throw in \n for linebreaks and flush manuall using std::flush once per frame or something) and for C you can use setvbuf to change the buffer mode on stdout, just make sure you have a big enough buffer for your needs.

If you are using debug builds quite a few third party libraries will push data through stdout or stderr to your IDE, make sure you are using release builds when you are profiling.

[size="1"]I don't suffer from insanity, I'm enjoying every minute of it.
The voices in my head may not be real, but they have some good ideas!
Servant of the Lord
Servant of the Lord

I'm still wondering, why fps would be useless number for profiling?

fps measures the entire loop, which means you only know total execution time.

Also, FPS is a deceptive measurement because it's not linear - the "cost" of each unit of FPS changes. Going from 59 to 60 FPS is not the same as going from 29 to 30 FPS.


Performance: Why did I just lose 100 FPS by drawing one triangle/sprite?

[Edit:] Lactose! already mentioned that. Oops. :ph34r:

Alberth
Alberth
EDIT: Guys, VerySleepy continues to show that printf takes 90-100% of some thread. I checked all parts of my source code, I even tried commenting everything, fps went crazy and the profiler still says that printf works in background somehow, although I checked multiple times that I don't have anything that says 'print' in my whole project.

The profiler cannot give you information where this "printf" function supposedly is located?

The only simple check I can think of now is to add


static bool once = true;
assert(once);
once = false;

to your 'snprintf' code, to really check you don't run it more than once.

(A breakpoint in that code would work too.)

alvaro
alvaro


void MMORPG::getBoneLocation(Shader& shader)
{
    for (unsigned int i = 0 ; i < 100 ; i++)
    {
        char Name[128];
        memset(Name, 0, sizeof(Name));
        snprintf(Name, sizeof(Name), "gBones[%d]", i);
        boneLocation[i] = glGetUniformLocation( shader.programID, Name );
    }
} 



I can think of several problems with that piece of code, but I'll point out two:
* You don't need to call memset at all. Just erase that line.
* getBoneLocation is a pretty bad name for a function that doesn't get you the location of a bone. Perhaps computeBoneLocations or precomputeBoneLocations would do.
Heelp
Heelp
The profiler cannot give you information where this "printf" function supposedly is located?

I would guess it's because the profiler was free and it doesn't have any advanced features. Maybe I have to download another one. I would ask you what is yours, but I know you don't use Windows.

I think SimonForsman is right, third-party libs mess up with the stdout, because I set up another project without SDL and openGL and there was no prinf() complaints from the profiler. I will leave this for now, because I don't want to touch compiler settings and mess things up.

getBoneLocation is a pretty bad name for a function that doesn't get you the location of a bone. Perhaps computeBoneLocations or precomputeBoneLocations would do.

Alvaro, but this is kind of what it does, because it gets the location of my gBones[] array in the vertex shader. But it doesn't really matter, I just copy/pasted it from one tutorial, haven't really thought about that.

Guys, last question: I'm wondering if it's ok to precompute all the matrix interpolations, this way I will leave the CPU with much less per-frame calculations. I doubt that it will take so much ram to store the matrices, but still I wanted to ask in case someone have tried that before?

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.