Skip to main content
GameDev.net gamedev.net
🔒 Locked

IOCP+AcceptEx vs blocking accept()

Started by Evil Steve May 20, 2009 at 9:19 AM 18 replies 11k views
Original Post
Evil Steve
Evil Steve
Hi all, I'm looking into IO Completion Ports for some test code I'm developing, and I have a quick question. I've been reading This Page (I know it's a bit dated), and it seems to imply that using AcceptEx is the correct way to accept connections when using IOCPs. However, I don't see any advantage to just using a blocking accept() in a thread without using IOCP at all. The only advantage I can see to using IOCPs for accepting connections is if the accept process takes a long time, so the server is spending a long time not waiting in an accept() call. So - am I missing something that makes IOCPs useful for accepting connections? Cheers, Steve
Antheus
Antheus
IOCP allows you to have outstanding operations. Simply have 50 accept calls pending on a socket. Then, if one of them takes long time, others will take over as long as there are workers that can process them.

This is the core principle behind IOCP. There can be arbitrary number of pending operations waiting for network events (read/write/accept), and their handlers will be dispatched as needed.

With synchronous model you always need to process one batch before starting with another. While simpler in implementation, misbehaving operations have more impact.
Evil Steve
Evil Steve
I see. So does that mean that it's a good idea to create a "lot" of threads (I.e. 8 or 16) to process AcceptEx requests for when connection frequency spikes? In theory the only down side would be a few redundant (blocked) threads when usage is low.

The code I'm writing is designed to be the network code for a large-ish multiplayer game or a web server - something designed to handle frequent connections and disconnections.

Thanks for the reply,
Steve
Antheus
Antheus
Quote:
Original post by Evil Steve
I see. So does that mean that it's a good idea to create a "lot" of threads (I.e. 8 or 16) to process AcceptEx requests for when connection frequency spikes? In theory the only down side would be a few redundant (blocked) threads when usage is low.


IOCP has its own thread pool which takes care of that, you merely need to have enough outstanding requests.

Application can use several threads if it can use them to do more work. But if each connection needs to lock the database, then having more than one application thread will not bring any benefit. IOCP will still be handling n accepts in different threads, but application would be unable to do more work.

As an example:
on_receive() {  query = parse_received_string();  database.lock();  result = database.query();  database.unlock();  client.send(result);}


Running these handlers in multiple application threads would improve performance since despite database being shared, parsing and sending could be performed concurrently.

But both receive and send would be handled concurrently by IOCP the way it deems best.

Practical tests show that when there is no contention, the number of application threads that number of threads needed is same as number of cores +1. The +1 comes simply from various cases where a bit more work could be done. For two core, thread times on 2 core machine would be something like (97%, 97%, 4%), even if there are 10 threads in the pool. Overhead usually increases with more cores (consequently more threads).


If blocking operations are involved, or expensive operations need to run inside locks, then it will be up to you to ensure they don't cause problems. Even if you put 100 threads in your pool, all of those could choose to block.

Quote:
The code I'm writing is designed to be the network code for a large-ish multiplayer game or a web server - something designed to handle frequent connections and disconnections.


Web servers can usually be distributed quite effectively by keeping each request handler's logic as local as feasible. So your application would maintain a thread pool that would be handling connection callbacks from IOCP (threads calling GetQueuedCompletionStatus/Ex).

Each logical connection (accept, multiple reads, logic, multiple writes, close) would then be handled sequentially. Once one of them completes, next step is performed in completion handler. This is where IOCP comes in - even though this is strictly sequential operation (and can be considered as single-threaded), it will multiplex hundreds of such connections doing the same in some predetermined number of threads. These will likely match the number of cores, since they will either be running at full load, or there will not be enough work to do. This is one part of equation.

But since logic may involve long running operations, blocking calls or large operations on shared state those need to be addressed separately. If each of requests needs to modify some shared counter or container, IOCP cannot help with that, since it's up to you to lock or otherwise handle the scenario.
Evil Steve
Evil Steve
Quote:
Original post by Antheus
Quote:
Original post by Evil Steve
I see. So does that mean that it's a good idea to create a "lot" of threads (I.e. 8 or 16) to process AcceptEx requests for when connection frequency spikes? In theory the only down side would be a few redundant (blocked) threads when usage is low.


IOCP has its own thread pool which takes care of that, you merely need to have enough outstanding requests.
Do I not still have to have multiple threads each calling AcceptEx()?

Quote:
Original post by Antheus
Application can use several threads if it can use them to do more work. But if each connection needs to lock the database, then having more than one application thread will not bring any benefit. IOCP will still be handling n accepts in different threads, but application would be unable to do more work.

[snip]
Yeah, I'm aware of what happens when a worker thread enters a blocking operation, and I've been pouring over the MSDN for various IOCP-related information.

Quote:
Original post by Antheus
Quote:
The code I'm writing is designed to be the network code for a large-ish multiplayer game or a web server - something designed to handle frequent connections and disconnections.


Web servers can usually be distributed quite effectively by keeping each request handler's logic as local as feasible. So your application would maintain a thread pool that would be handling connection callbacks from IOCP (threads calling GetQueuedCompletionStatus/Ex).
I'd like to keep things as generic as possible here (I know that generic + high performance doesn't really go together) - for the multiplayer game aspect, I'd be using the accepting threads to accept the connection, create a new player information struct, and shove it in a queue of logging in or connected players, meaning the accepting threads wouldn't be doing a huge amount of activity, and in theory won't be blocking (Using lockless data structures and no heap access where possible).

Cheers,
Steve
aissp
aissp
Advantage is very clear - thread context switching. Blocking accept will awake thread and cause performance degradation.

Here is very useful ref about production server design:

http://www.kegel.com/c10k.html

M$'ve (with their iocp) just implemented one of the possible strategy...
hplus0603
hplus0603
With IOCP, you issue requests from any thread, and they complete on one of the thread pool threads. Thus, you can issue a dozen AcceptEx requests all from the main thread, and then when an AcceptEx completes in the thread pool, you re-issue one AcceptEx as part of the completion handling.

However, I don't see why you'd want more than one or perhaps two outstanding AcceptEx requests. If you're so overloaded on requests that calling AcceptEx takes a long time (looking up an IP address in a hash table of blocked addresses, perhaps?) then you wouldn't have the server resources to actually serve the users when they're accepted anyway.
enum Bool { True, False, FileNotFound };
Evil Steve
Evil Steve
Quote:
Original post by hplus0603
With IOCP, you issue requests from any thread, and they complete on one of the thread pool threads. Thus, you can issue a dozen AcceptEx requests all from the main thread, and then when an AcceptEx completes in the thread pool, you re-issue one AcceptEx as part of the completion handling.
So is it usually preferable to have two IOCPs then, one for handling send/recv from clients and one for handling accept requests from new sockets?
I was going to have two accepting worker threads, each one calling AcceptEx(), then GetQueuedCompletionStatus() and then processing the new connection in a loop. Should I instead make the AcceptEx() call from the main thread for whatever reason? Or does it really not matter either way?

Quote:
Original post by hplus0603
However, I don't see why you'd want more than one or perhaps two outstanding AcceptEx requests. If you're so overloaded on requests that calling AcceptEx takes a long time (looking up an IP address in a hash table of blocked addresses, perhaps?) then you wouldn't have the server resources to actually serve the users when they're accepted anyway.
That's a good point, yeah.

Cheers,
Steve
Erik Rufelt
Erik Rufelt
You seem to misunderstand IOCP, have you used it before?
Only ever have one completion port. It will handle everything, from WriteFile to WSASend and AcceptEx. The point is that the OS handles everything you post behind the scenes, in a way we are to assume optimal, and a completion notification will be returned and handled from GetQueuedCompletionStatus whenever an operation has completed or failed. Then there are some number of threads waiting on GetQueuedCompletionStatus, each thread exactly identical, and the OS will only ever allow exactly the same amount of threads to run at the same time as there are processors on the system. Theoretically this gives optimal performance, or at least that's the idea.
I think the best about it is that once a software is designed properly to use it, it's a very beautiful and simple system, yet very powerful.
Antheus
Antheus
The idea behind asynchronous connection-centric networking (such as HTTP) is that it creates a sort of cooperative multi-threading model. At conceptual level, the code looks like this (all function calls are non-blocking):
on_start() {  accept(on_accept())}on_accept() {  start()        // start listening again  read(on_read)}on_read() {  if (read_enough) {    process(on_processed)  } else {    read(on_read)  }}on_processed() {  send(on_send)}on_send() {  if (done_sending)    close()  else send(on_send)}


This is equivalent code to:
accept()while (!read_enough) read()process()while (!done_sending) send()close()


Since with this model each connection has only one outstanding request at any given time, it behaves as if everything were single-threaded, even without locking, even if there are many connections active at the same time.
Evil Steve
Evil Steve
I see - but why would I not want to have two IOCPs? Surely that'd allow me to have two pools of threads; one for each IOCP, and each thread can be concerned with either accepting connections or send/recv calls?
I don't see how it'd be less efficient that way.
hplus0603
hplus0603
Why would you want to do accept processing from another set of threads than your other I/O? What do you expect to win by doing that?

In general, no, two sets of threads is a bad idea, because it puts more pressure on the kernel, allocates more stack space, and doesn't allow you to make the best use of available cores. If there are 4 CPU cores, it doesn't make sense to have more than 4 threads actually running at the same time. With a single I/O completion port for all asynchronous requests, you can regulate it to be optimal. If you have multiple separate sets of threads, that kind of regulation is either impossible, or a lot harder, for no real benefit.

Note that Windows has a thread pool mechanism built-in, which may be better at regulating thread count than you can do on your own. In XP and earlier OS-es, that mechanism is somewhat limited and hard to work with; in Vista and Server 2008 and up, there is a new thread pool API that's a lot better to work with.

But... why not use boost::asio and let someone else sweat the small stuff? :-)
enum Bool { True, False, FileNotFound };
Evil Steve
Evil Steve
Quote:
Original post by hplus0603
Why would you want to do accept processing from another set of threads than your other I/O? What do you expect to win by doing that?

In general, no, two sets of threads is a bad idea, because it puts more pressure on the kernel, allocates more stack space, and doesn't allow you to make the best use of available cores. If there are 4 CPU cores, it doesn't make sense to have more than 4 threads actually running at the same time. With a single I/O completion port for all asynchronous requests, you can regulate it to be optimal. If you have multiple separate sets of threads, that kind of regulation is either impossible, or a lot harder, for no real benefit.
I see, fair enough. The suggestion for two sets of completion ports was just to make the code a bit tidier - the worker threads don't need to deal with both accept and send/recv requests, but it's not really that much of a win.

Quote:
Original post by hplus0603
Note that Windows has a thread pool mechanism built-in, which may be better at regulating thread count than you can do on your own. In XP and earlier OS-es, that mechanism is somewhat limited and hard to work with; in Vista and Server 2008 and up, there is a new thread pool API that's a lot better to work with.

But... why not use boost::asio and let someone else sweat the small stuff? :-)
Well, my test server is XP-64, which is a bit old and crappy, so I think I'll stick with my own thread pool for now (Static sized, I seem to recall that the OS thread pool resizes and does other fun things).
My main reason for not using boost:asio or another library is for the learning experience - I've attempted IOCP a couple of times before and given up pretty quickly, or ended up with some horrible mess of critical sections which constantly deadlocks, crashes, or just performs poorly.
Erik Rufelt
Erik Rufelt
Unless you have insane amounts of traffic the difference in efficiency will be minimal, but as hplus0603 said there's no reason to. Just put a switch statement for the type of operation completed in the thread, and call the appropriate function, or create a base-class for objects that handle an operation. It should give cleaner code and make it very easy to add new types of operations too, without the need for creating new threads or something.
Evil Steve
Evil Steve
Quote:
Original post by Erik Rufelt
Unless you have insane amounts of traffic the difference in efficiency will be minimal, but as hplus0603 said there's no reason to. Just put a switch statement for the type of operation completed in the thread, and call the appropriate function, or create a base-class for objects that handle an operation. It should give cleaner code and make it very easy to add new types of operations too, without the need for creating new threads or something.
Yeah, I'll just shove a switch() in there - There'll need to be a check for read vs write operations anyway.

Ok, I think I'm all set for now - thanks for all the replies!
hplus0603
hplus0603
Typically what you do is set the user ref pointer (or other similar reference) to a pointer to a completion interface, and pass in the appropriate object type for the operation. Thus, calling obj->onComplete() will do read, write or accept depending on what operation the user requested. No switch() necessary.

In C++, whenever you find yourself using switch(), you should ask yourself whether you can use virtual functions instead, and you'll generally end up with a cleaner implementation/interface.
enum Bool { True, False, FileNotFound };
aissp
aissp
Oh, to clear the picture a little:
(1) When we create listen socket, the OS create two lists which assigned with this socket - incomplete connection queue (ICU) and pending connection queue (PCU).
(2) When OS get SYN request it's create "socket" structure and put it in ICU.
(3) When handshake is totally finished OS put the socket in PCU and scan if there are any accept function waiting, if yes - this accept function will call.

So, accept call only for totally connected socket.
(3a) if socket in PCU receive any data, this data will be stored in input buffer of this socket. So, before accept calling the TCP connection exists and complete functional (one of the advantage AcceptEx it can read data with accept connection).

If the list PCU is full the client with send SYN will not get any answer from server (as we suggest that we can clear the PCU fast enough)

Some digits: the socket will exist in PCU without accept a significant time (i guess it should be about 2 minutes for linux systems). What is a typical size for PCU? Oracle recommend for its http server 1024 for ICU and 128 for PCU. My opinion, it is large enough for online games, so i'm happy with 10 for PCU, and only one AcceptEx for iocp thread pool.

Hope this help
Antheus
Antheus
Quote:
Original post by Evil Steve

Yeah, I'll just shove a switch() in there - There'll need to be a check for read vs write operations anyway.


If you can stomach the bloated syntax, look at asio tutorials. They are not immediately transferable to IOCP, but they show the general structure of asynchronous design.

Even if portability is not a requirement, asio has the benefit of providing some facilities which you would need to reinvent anyway:
- timers (for timeouts) implemented in a fairly optimal manner
- control over how handlers are dispatched (strands)
- generic way to pass specific handlers for each operation (via bind, allowing support for arbitrary parameters)
- one-line support for multiple application threads
- allows transparent use of custom asynchronous logic not tied to CP (io_service::dispatch() and io_service::post())

The downsides are long compile times and gnarly syntax.
hplus0603
hplus0603
Quote:
Original post by aissp
Some digits: the socket will exist in PCU without accept a significant time (i guess it should be about 2 minutes for linux systems).


Actually, on Linux with SYNcookies, no entry will be put into any list until the ACK comes back from the SYNACK that the server sends. This allows the kernel to avoid allocating any data for certain kinds of DOS attacks (SYN floods etc).

I think the main point is important, though: In networking, the kernel will queue and buffer almost everything for you. The sockets API is a way of getting data to and from the kernel, not to and from the network device. Thus, the kind of shenanigans you have to keep up when you write device drivers or deal with double-buffered media, are not necessary or helpful in most networking APIs.
enum Bool { True, False, FileNotFound };
Fallonsnest
Fallonsnest
I have a strange question about the completion port.

I think most servers run as console application, so also have to deal with console input/output. I dont know how threadsafe method like WriteConsole are as the SDK doc's dont mention it.

So what about async console/file handling in a iocp socket application?
Should the console and file handles be bound to the same completion port or should there be a second completion port for handling those (with its own worker threads)

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.