Original Post
Hi all! First of all, please excuse any language errors - as English is not my native language. Since our current game is in an art-heavy stall - I'm able to really take the time to design our next game engine from scratch - with a (at the moment) very narrow planned user base (Windows Vista+, DirectX 10+ and most likely multi-core processors under the hood) - with no planned support for downscaling to lesser cards or sidestepping into other Operating Systems apart from later incarnations of Windows. (these facts might influence the issues presented below) In order to fully utilize the potential of the system at hand I will have to utilize the amount of available cores at our disposal as well as possible. Now, the initial idea was to start threading the stuff that could be run in parallel as usual - but I was thinking about if another approach would potentially lead to better parallelism and scale better with potential future many-core processors. This is where I feel I'm not fully aware of the implications, and thus seek your knowledge. The other approach consists basically of a host "kernel process" and separate worker processes (not threads but actual processes) that each have predefined tasks. The worker processes are to be buffered - and use shared memory (through memory mapped files) and use some sort of compare-and-swap to handle messaging between the processes. This would imply a slight latency between what is rendered and what the world state is in (i.e. what's rendered isn't the whole truth) - but hopefully I can arrange that latency to hit certain tasks harder than others. (keeping input and local actor highly up to date while allowing a bigger latency for less important actors) Now, it might seem such an approach is overly complicating things - but it would have a few nice perks to go along with it. First of all, it would force us to think about parallelism at all times, since we can never be sure in what order things will occur without a sync-lock from the kernel process. It feels like we'd automatically have less locks on data due to the separate application pool and buffered nature of the system and it gives a rather nice way of handling the update rate of specific tasks. We'd also potentially be able to detect crashes of separate subsystems from the kernel and try to handle those gracefully. Now - before I get to involved into the design of the second approach, I have to ask: Am I shooting myself in the foot here? Would performance plunge by utilizing separate processes with shared memory to handle the interchange of data? (only interested in the performance side of things, not the manageability or debug-friendliness of the solution) I'm fully aware that buffering the data will imply a higher memory cost - but regardless of that fact and assuming I could pull it off - would (could) it work at least near a standard multi-threading solution, performance-wise?