Skip to main content
GameDev.net gamedev.net
🔒 Locked

Multi-processes rather than/in addition to multi-threading?

Started by DEVLiN Sep 6, 2008 at 3:37 PM 11 replies 2.9k views
Original Post
DEVLiN
DEVLiN
Hi all! First of all, please excuse any language errors - as English is not my native language. Since our current game is in an art-heavy stall - I'm able to really take the time to design our next game engine from scratch - with a (at the moment) very narrow planned user base (Windows Vista+, DirectX 10+ and most likely multi-core processors under the hood) - with no planned support for downscaling to lesser cards or sidestepping into other Operating Systems apart from later incarnations of Windows. (these facts might influence the issues presented below) In order to fully utilize the potential of the system at hand I will have to utilize the amount of available cores at our disposal as well as possible. Now, the initial idea was to start threading the stuff that could be run in parallel as usual - but I was thinking about if another approach would potentially lead to better parallelism and scale better with potential future many-core processors. This is where I feel I'm not fully aware of the implications, and thus seek your knowledge. The other approach consists basically of a host "kernel process" and separate worker processes (not threads but actual processes) that each have predefined tasks. The worker processes are to be buffered - and use shared memory (through memory mapped files) and use some sort of compare-and-swap to handle messaging between the processes. This would imply a slight latency between what is rendered and what the world state is in (i.e. what's rendered isn't the whole truth) - but hopefully I can arrange that latency to hit certain tasks harder than others. (keeping input and local actor highly up to date while allowing a bigger latency for less important actors) Now, it might seem such an approach is overly complicating things - but it would have a few nice perks to go along with it. First of all, it would force us to think about parallelism at all times, since we can never be sure in what order things will occur without a sync-lock from the kernel process. It feels like we'd automatically have less locks on data due to the separate application pool and buffered nature of the system and it gives a rather nice way of handling the update rate of specific tasks. We'd also potentially be able to detect crashes of separate subsystems from the kernel and try to handle those gracefully. Now - before I get to involved into the design of the second approach, I have to ask: Am I shooting myself in the foot here? Would performance plunge by utilizing separate processes with shared memory to handle the interchange of data? (only interested in the performance side of things, not the manageability or debug-friendliness of the solution) I'm fully aware that buffering the data will imply a higher memory cost - but regardless of that fact and assuming I could pull it off - would (could) it work at least near a standard multi-threading solution, performance-wise?
AndyPandyV2
AndyPandyV2
Why? What possible benefit would there be to doing this instead of multiple threads?
Anonymous P
Anonymous P
Quote:
The worker processes are to be buffered - and use shared memory (through memory mapped files) and use some sort of compare-and-swap to handle messaging between the processes.
So you want a homebake message passing system; you want the slowness of shared-nothing message passing combined with the errors and complexity of manual shared memory synchronization and lockfree structures. Sounds like a win.
Quote:
Now, it might seem such an approach is overly complicating things
To say the least.
Quote:
First of all, it would force us to think about parallelism at all times, since we can never be sure in what order things will occur without a sync-lock from the kernel process.
You can't be sure of what order things will occur in different threads without a sync primitive serializing them either: they force you to think about parallelism in exactly the same way...
Quote:
It feels like we'd automatically have less locks on data due to the separate application pool and buffered nature of the system and it gives a rather nice way of handling the update rate of specific tasks.
So use sockets and do without shared memory. You're getting the worst of both worlds by writing homebrew sockets on top of shared memory.
Quote:
Now - before I get to involved into the design of the second approach, I have to ask: Am I shooting myself in the foot here?
You seem to want to implement sockets on top of shared memory; a great way to combine the drawbacks of both methods into one buggy and inefficient mess.
Quote:
Would performance plunge by utilizing separate processes with shared memory to handle the interchange of data? (only interested in the performance side of things, not the manageability or debug-friendliness of the solution)
Depends completely on your problem. If you have a big resource that is read often and updated, threads will beat the crap out of shared nothing message passing.

Quote:

I'm fully aware that buffering the data will imply a higher memory cost - but regardless of that fact and assuming I could pull it off - would (could) it work at least near a standard multi-threading solution, performance-wise?
I'd say it'd work near a 'solution' using sockets, performance-wise, and take a lot more work.
DEVLiN
DEVLiN
If I somehow offended anyone I apologize, but I can't find anything in my original post that sparked such hostility in the replies.

I'm also not sure why you brought sockets into this. In quite a few posts on these very forums (a simple search for: shared memory sockets, for instance) - people recommend shared memory over sockets for inter-process communication for precisely performance reasons. The lists I've seen imply that shared memory using semaphores is the fastest and sockets the slowest way of inter-process communication. Quite different from what your reply suggested.

I'll rephrase my question to something more general instead:

Is there any inherent, large difference performance wise between using multiple threads and multiple processes - or is it solely dependent on how it's used?
abdulla
abdulla
The biggest overhead you'll have by going multi-process is context switching. It's a lot faster to switch between threads in the same address space. And you're correct, shared memory is faster than sockets due in part to reduction in the amount of syscalls needed. I'd also advise against file-backed shared memory, I'm not sure how it's done under Windows, but there's shm_open and related functions in POSIX that allow you to create and manage shared memory without a file backing.

All in all, unless you need the added protection of address space separation, you're better off sticking with a multi-threaded design.
swiftcoder
swiftcoder
Quote:
Original post by abdulla
And you're correct, shared memory is faster than sockets due in part to reduction in the amount of syscalls needed.
Only if you are able to avoid locking - as soon as you have to lock your shared memory, the sockets do about as well, since a local socket (at least in the unix world) tends to be implemented as chunk of shared memory and some locking primitive, which evens out your syscalls again.
Tristam MacDonald. Ex-BigTech Software Engineer. Future farmer. [https://trist.am]
abdulla
abdulla
Quote:
Original post by swiftcoder
Quote:
Original post by abdulla
And you're correct, shared memory is faster than sockets due in part to reduction in the amount of syscalls needed.
Only if you are able to avoid locking - as soon as you have to lock your shared memory, the sockets do about as well, since a local socket (at least in the unix world) tends to be implemented as chunk of shared memory and some locking primitive, which evens out your syscalls again.


Actually, I've got some benchmarks for a paper I'm working on that show that that's a false assumption. You end up copying around data a lot more, even if you use tricks like vmsplice (which require you to allocate pages of memory). The overhead for semaphores, especially under Linux, is quite low, and you can use spin locks to reduce the latency further.
swiftcoder
swiftcoder
Quote:
Original post by abdulla
Quote:
Original post by swiftcoder
Quote:
Original post by abdulla
And you're correct, shared memory is faster than sockets due in part to reduction in the amount of syscalls needed.
Only if you are able to avoid locking - as soon as you have to lock your shared memory, the sockets do about as well, since a local socket (at least in the unix world) tends to be implemented as chunk of shared memory and some locking primitive, which evens out your syscalls again.
Actually, I've got some benchmarks for a paper I'm working on that show that that's a false assumption. You end up copying around data a lot more, even if you use tricks like vmsplice (which require you to allocate pages of memory). The overhead for semaphores, especially under Linux, is quite low, and you can use spin locks to reduce the latency further.
I didn't really state my point very well there, so here goes:
a) syscall count is not a good metric of performance.
b) context switches are expensive when working with processes - when you start blocking you are going to lose.
c) you are most likely going to need to do those extra copies to make your shared-memory solution non-bocking.
Tristam MacDonald. Ex-BigTech Software Engineer. Future farmer. [https://trist.am]
abdulla
abdulla
Quote:
Original post by swiftcoder
I didn't really state my point very well there, so here goes:
a) syscall count is not a good metric of performance.
b) context switches are expensive when working with processes - when you start blocking you are going to lose.
c) you are most likely going to need to do those extra copies to make your shared-memory solution non-bocking.


a) No it's not, but it has a big effect on performance. Reducing syscalls really does help, but this is at the nanosecond level (at least on my test systems).
b+c) I guess it really comes down to whether you want to block or not, so it depends on the structure of the system.
swiftcoder
swiftcoder
Quote:
Original post by abdulla
b+c) I guess it really comes down to whether you want to block or not, so it depends on the structure of the system.
If you are using threads, blocking isn't a big deal, but with processes you need to be non-blocking, to avoid that painful context switch.

The only area where processes have an advantage over threads is stability (one process crashing shouldn't affect another), however since in your original scenario the processes are interdependent, you wont benefit from this either.
Tristam MacDonald. Ex-BigTech Software Engineer. Future farmer. [https://trist.am]
Antheus
Antheus
Quote:
The other approach consists basically of a host "kernel process" and separate worker processes (not threads but actual processes) that each have predefined tasks.


There are several criteria you can use to judge whether a design is suitable or not. Process, thread, network separation are all implementation details.

One criteria is responsiveness and graceful degradation. For complex calculation which do not require responsiveness the system must be designed to complete given tasks, taking into consideration stalls. If responsiveness is at premium, then at some point, system will need to "crash".

Which to prioritize depends, for real-time systems responsiveness is typically most important and enforced by hard caps (n users, m actions, o resources, ... at most).

Enterprise systems typically prefer graceful degradation at expense of larger constant factor costs. In many cases, waiting several minutes for a guaranteed result of a query is preferable to having system respond instantly, but failing if overloaded.

Shared memory, pipes and other forms of IPC are perfectly valid. All of them can be tailored to meet performance characteristics (example).


For responsive and scalable systems that offer some leverage for degradation, task-based parallelism over mostly read-only shared state is generally preferred. This approach makes threads more convenient in general due to lower resource use and potentially perfect CPU utilization. It's also most suitable for coarse-grained concurrency.


Process-per-task is sometimes used for shared-nothing loosely coupled systems. The by far biggest advantage is redundancy - but redundancy is a moot point if such system runs locally only. If it is important, than separate process approach offers some advantages.


In general however, tasks that need to operate on same shared data should be close. Typical example is collision detection. While at first it may seem like a good idea to perform various distributed queries among processes, constant factor introduced by communication overhead as well as replication of state and global resources becomes prohibitively high. (consider distributed collision detection over size of earth, where on each tick you need to synchronize billions of entities - clearly, each process needs entire geometry - so it might as well run everything within a single process).

Process-per-task is generally best suited for completely de-coupled processes which can be reduced to black-box (input/output) function. An extreme example of this is MapReduce, which leverages redundancy with arbitrary scalability - it is not however useful for anything resembling real-time.

Quote:
Would performance plunge by utilizing separate processes with shared memory to handle the interchange of data? (only interested in the performance side of things, not the manageability or debug-friendliness of the solution)


The only thing that will determine the "performance" is how resources and tasks are distributed. This approach can be optimal or even superior, but by itself, there are no guarantees.

Solving the problem of concurrency however is mostly unrelated to words. Process or thread are minor implementation issues which by themselves have no impact on results. The choice will almost always depend exclusively on other factors (deployment, administration, redundancy).

Quote:
but I was thinking about if another approach would potentially lead to better parallelism and scale better with potential future many-core processors


Only application design will determine that. Pros/cons of using either are mostly irrelevant, neither of them solves the problem of transparent scaling. (threads are generally slightly simpler if there's proper task scheduling system in place).

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.