Skip to main content
GameDev.net gamedev.net
🔒 Locked

switch or if, whats faster?

Started by pavel989 Sep 10, 2008 at 12:39 AM 15 replies 2.8k views
Original Post
pavel989
pavel989
execution wise, are they the same, or is one faster than the other?
hh10k
hh10k
I would say that it's within the power of a compiler to make code just as fast using either, however it's easier for the compiler to apply certain optimisations to switch statements. Any decent compiler should turn large switch statements into a jump table (when it is suitable).
pinacolada
pinacolada
I would have guessed that the if statement is faster, since it can usually be represented by one "jump if" operand.

But anyway, they are both fast enough. Asking if one is faster is like a car designer asking, "what color of paint makes my car go fastest?"
pavel989
pavel989
ah i see, ty. im just trying to figure out, why within my program, there is a slow process at one point. (its an opengl program)

im using the HWND stuff, and all its really doing is moving a point as the pointer moves, and the switch statement was simply telling it which point to deal with.

if the mouse flies out the window really fast, the points stop far behind the edge, so i thought the switch was slowing it down.

hm, gotta keep checking stuff.

again ty
Chaotic_Attractor
Chaotic_Attractor
The truth is that with today's processors such comparisons are really irrelevant. So what if one uses a few more clock cycles than the other? If the switch statement makes my code more readable that's what I'm going to use, unless you're working on a project where every nanosecond counts (which I'm certain you aren't).
Strength without justice is no vice; justice without strength is no virtue.
Mxz
Mxz
If you take Visual Studio 2005 as an example, and the following two code samples as a simple test.

Switch statement version
int main(){	int i = rand() % 5;	switch(i)	{	case 0:	i = 0;		break;	case 1:	i = 1;		break;	case 2:	i = 2;		break;	case 3:	i = 3;		break;	case 4:	i = 4;		break;	case 5:	i = 5;		break;	}	std::cout << i;}



if-else statement version
int main(){	int i = rand() % 5;	if(i == 0)	{		i = 0;	}	else if(i == 1)	{		i = 1;	}	else if(i == 2)	{		i = 2;	}	else if(i == 3)	{		i = 3;	}	else if(i == 4)	{		i = 4;	}	else if(i == 5)	{		i = 5;	}	std::cout << i;}


Under the compiler optimization setting Favor fast code, the following disassembly is produced for the switch statement version.

   call        dword ptr [__imp__rand (4020A8h)]   cdq                mov         ecx,5   idiv        eax,ecx   cmp         edx,ecx   ja          $LN1+2 (401081h)   jmp         dword ptr  (401094h)[edx*4] $LN6:  mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   xor         edx,edx   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret              $LN5:  mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   mov         edx,1   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret              $LN4:  mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   mov         edx,2   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret              $LN3:  mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   mov         edx,3   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret              $LN2:  mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   mov         edx,4   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret              $LN1:  mov         edx,ecx   mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret                lea         ecx,[ecx] 


As you can see, there are just a couple of jumps at the top of the function. However, the same code under the Favor small code optimization setting compiles down to the following.

  call        dword ptr [__imp__rand (4020A8h)]   cdq                push        5      pop         ecx    idiv        eax,ecx   mov         eax,edx   sub         eax,0   je          main+37h (401037h)   dec         eax    je          main+32h (401032h)   dec         eax    je          main+2Dh (40102Dh)   dec         eax    je          main+29h (401029h)   dec         eax    je          main+25h (401025h)   dec         eax    jne         main+39h (401039h)   push        ecx    jmp         main+2Fh (40102Fh)   push        4      jmp         main+2Fh (40102Fh)   push        3      jmp         main+2Fh (40102Fh)   push        2      pop         edx    jmp         main+39h (401039h)   xor         edx,edx   inc         edx    jmp         main+39h (401039h)   xor         edx,edx   mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret  


On this compiler setting the select statement has been turned into a series of comparisons and jumps.

The if statements on the other hand produce the following output when the Favor fast code optimization is set.

  call        dword ptr [__imp__rand (4020A8h)]   cdq                mov         ecx,5   idiv        eax,ecx   test        edx,edx   jne         main+22h (401022h)   mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret                cmp         edx,1   jne         main+37h (401037h)   mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret                cmp         edx,2   jne         main+4Ch (40104Ch)   mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret                cmp         edx,3   jne         main+61h (401061h)   mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret                cmp         edx,4   jne         main+76h (401076h)   mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret                cmp         edx,ecx   jne         main+7Ch (40107Ch)   mov         edx,ecx   mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret   


Here we have many comparisons and jumps taking place.

The Favor small code setting produces the following out of the if statement.

  call        dword ptr [__imp__rand (4020A8h)]   push        5      cdq                pop         ecx    idiv        eax,ecx   test        edx,edx   je          main+34h (401034h)   cmp         edx,1   je          main+34h (401034h)   cmp         edx,2   jne         main+1Dh (40101Dh)   push        edx    jmp         main+33h (401033h)   cmp         edx,3   jne         main+25h (401025h)   push        edx    jmp         main+33h (401033h)   cmp         edx,4   jne         main+2Dh (40102Dh)   push        edx    jmp         main+33h (401033h)   cmp         edx,5   jne         main+34h (401034h)   push        edx    pop         edx    mov         ecx,dword ptr [__imp_std::cout (40203Ch)]   push        edx    call        dword ptr [__imp_std::basic_ostream<char,std::char_traits<char> >::operator<< (402038h)]   xor         eax,eax   ret  


As you can see, for Visual Studio the if statements in this case have been implemented using many jumps under both compiler settings. The select statement was implemented using a jump table when using the favor fast code optimization setting, but was implemented using a series of comparisons and jump statements under the favor small code setting.

In short it depends on your compiler, code and compiler settings. It is hard to make general statements about these things. The best way is to profile and see.

[Edited by - Mxz on September 10, 2008 1:32:36 AM]
PrettyBoyTim
PrettyBoyTim
I'm pretty sure you're looking for the optimisation in the wrong place.

From the sounds of it, your problem is to do with the rate that you're getting input messages into your application. When you move the mouse fast, there will be a larger difference between subsequent mouse positions that are sent to your program. Therefore if you move the mouse fast out of your program's window, the last mouse position message it gets will be a fair distance from the side of the window. If you wish to fix this you will need to intercept mouse move messages that happen outside your window as well.
Evil Steve
Evil Steve
Quote:
Original post by Mxz
If you take Visual Studio 2005 as an example, and the following two code samples as a simple test.

Switch statement version
*** Source Snippet Removed ***


if-else statement version
*** Source Snippet Removed ***

Under the compiler optimization setting Favor fast code, the following disassembly is produced for the switch statement version.

*** Source Snippet Removed ***

As you can see, there are just a couple of jumps at the top of the function. However, the same code under the Favor small code optimization setting compiles down to the following.

*** Source Snippet Removed ***

On this compiler setting the select statement has been turned into a series of comparisons and jumps.

The if statements on the other hand produce the following output when the Favor fast code optimization is set.

*** Source Snippet Removed ***

Here we have many comparisons and jumps taking place.

The Favor small code setting produces the following out of the if statement.

*** Source Snippet Removed ***

As you can see, for Visual Studio the if statements in this case have been implemented using many jumps under both compiler settings. The select statement was implemented using a jump table when using the favor fast code optimization setting, but was implemented using a series of comparisons and jump statements under the favor small code setting.

In short it depends on your compiler, code and compiler settings. It is hard to make general statements about these things. The best way is to profile and see.
Is that in a release build? I would have thought the compiler would optimise out the entire if / switch statement in this case...
pavel989
pavel989
Quote:
Original post by PrettyBoyTim
I'm pretty sure you're looking for the optimisation in the wrong place.

From the sounds of it, your problem is to do with the rate that you're getting input messages into your application. When you move the mouse fast, there will be a larger difference between subsequent mouse positions that are sent to your program. Therefore if you move the mouse fast out of your program's window, the last mouse position message it gets will be a fair distance from the side of the window. If you wish to fix this you will need to intercept mouse move messages that happen outside your window as well.


well i just read a few hours ago, the msg thing, its a queue, so my program should be getting msgs down to the last one in the list, i just think its not sending data fast enof. who knows, ima look into the msg system and see what alternative or improvements there are.
Evil Steve
Evil Steve
Quote:
Original post by pavel989
Quote:
Original post by PrettyBoyTim
I'm pretty sure you're looking for the optimisation in the wrong place.

From the sounds of it, your problem is to do with the rate that you're getting input messages into your application. When you move the mouse fast, there will be a larger difference between subsequent mouse positions that are sent to your program. Therefore if you move the mouse fast out of your program's window, the last mouse position message it gets will be a fair distance from the side of the window. If you wish to fix this you will need to intercept mouse move messages that happen outside your window as well.


well i just read a few hours ago, the msg thing, its a queue, so my program should be getting msgs down to the last one in the list, i just think its not sending data fast enof. who knows, ima look into the msg system and see what alternative or improvements there are.
The message pump shouldn't be a bottleneck at all. Considering your CPU can execute hundreds of millions of jumps per second, a switch() versus and if() will make no relevant difference.

Quote:
Original post by pavel989
if the mouse flies out the window really fast, the points stop far behind the edge, so i thought the switch was slowing it down.
That's because you don't get WM_MOUSEMOVE messages when the pointer is outside of the window. If you need that, use SetCapture. It doesn't matter how fast you process messages, you won't get one for every single pixel the mouse travels over.
Antheus
Antheus
Quote:
Original post by pavel989

well i just read a few hours ago, the msg thing, its a queue, so my program should be getting msgs down to the last one in the list, i just think its not sending data fast enof. who knows, ima look into the msg system and see what alternative or improvements there are.


Mouse updates position at 50-200Hz. These messages may be further filtered, so that you only receive the last position.

Even if you write your own driver, you will never receive updates for every pixel cursor travels over.

This is a problem of sampling.

Considering that mouse cursor can be moved across entire screen in a fraction of a second, one would need the updates to exceed thousands of Hz to obtain guaranteed per-pixel accuracy.
Deyja
Deyja
In the case of the example posted earlier, the smaller version with more jumps might actually run faster because it is smaller - it fits in the cache. Additionally, erroneous branch predictions can cause the large jump table in the switch version to stall, while the if-elseif chain, with nothing but small jumps, would run faster.
Daaark
Daaark
I've always thought of the switch statement as simple syntax sugar for what would be an overcomplicated to type and read series of if statement blocks.
Simian Man
Simian Man
Quote:
Original post by Deyja
Additionally, erroneous branch predictions can cause the large jump table in the switch version to stall, while the if-elseif chain, with nothing but small jumps, would run faster.


This is a good point. Branch predictors have more trouble with indirect branches than direct ones. The jump table optimization is only really worth it with a rather large number of cases.

Quote:
Original post by Daaark
I've always thought of the switch statement as simple syntax sugar for what would be an overcomplicated to type and read series of if statement blocks.


No. The reason you can only switch on integral types such as ints, chars or enums is that the switch is designed to be easily convertible into a jump table.
pavel989
pavel989
btw, mxz ty for that really in depth explanation.
Neverdone
Neverdone
This is from C++ optimizations on AMD's website :

#6. Use if-else statements in place of switch statements that have noncontiguous case expressions

If the case expressions are contiguous or nearly contiguous integer values, most compilers translate the switch statement as a jump table instead of a comparison chain. Jump tables generally improve performance because they reduce the number of branches to a single procedure call, and shrink the size of the control-flow code no matter how many cases there are. The amount of control-flow code that the processor must execute is also the same for all values of the switch expression.

However, if the case expressions are noncontiguous values, most compilers translate the switch statement as a comparison chain. Comparison chains are undesirable because they use dense sequences of conditional branches, which interfere with the processor's ability to successfully perform branch prediction. Also, the amount of control-flow code increases with the number of cases, and the amount of control-flow code that the processor must execute varies with the value of the switch expression.

For example, if the case expression are contiguous integers, a switch statement can provide good performance:

switch (grade)
{
case 'A':
...
break;
case 'B':
...
break;
case 'C':
...
break;
case 'D':
...
break;
case 'F':
...
break;

But if the case expression aren't contiguous, the compiler may likely translate the code into a comparison chain instead of a jump table, and that can be slow:

switch (a)
{
case 8:
// Sequence for a==8
break;
case 16:
// Sequence for a==16
break;
...
default:
// Default sequence
break;
}

In cases like this, replace the switch with a series of if-else statements:

if (a==8) {
// Sequence for a==8
}
else if (a==16) {
// Sequence for a==16
}
...
else {
// Default sequence
}


pavel989
pavel989
well that isnt the case, but definetly a good thing to take note of, ty. ive always figured u should try to stay consistent but i didnt itd matter.

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.