Skip to main content
GameDev.net gamedev.net
🔒 Locked

cudaMemcpy & DirectX11

Started by OctavianTheFirst Jan 6, 2011 at 9:24 AM 5 replies 3.1k views
Original Post
OctavianTheFirst
OctavianTheFirst
Hello,

I was wandering if there is an equivalent of cudaMemcpy for DirectX11.

http://developer.download.nvidia.com/compute/cuda/2_3/toolkit/docs/online/group__CUDART__MEMORY_g48efa06b81cc031b2aa6fdc2e9930741.html#g48efa06b81cc031b2aa6fdc2e9930741

The interesting thing about cudaMemcpy is that it can copy from Host to Device and Device to Device.

So my question is in fact is there a way to copy data from Device to Device without taking it to system memory? (And that works for ATI&Nvidia)

So instead of using AFR (Alternate Frame Rendering) or any pre-coded thing, the coder can decide what each GPU does, which seems a much better option to me!

Also, is it possible to de-activate the crossfire/sli link on dual-GPU single cards, and use them as 2 different devices? (And use the cudaMemcopy/equivalent to transfer from one to another?!)

Thanks!
DieterVW
DieterVW
DirectX has no API's for the proprietary hardware setups such as SLI and CrossFire nor is there anything for GPU to GPU transfers which would also be proprietary at this point.
DieterVW
DieterVW
Yes, CopyResource() lets you transfer data from one resource to another on the same device, but there is still no way to do transfers from GPU1 to GPU2 without first going through the CPU.
OctavianTheFirst
OctavianTheFirst
Is it because they're connected using the PCIE lanes directly to the CPU? Or are there some interconnect lanes?

I believe cudaMemcpy uses PCI-E (no SLI) to transfer data directly from a GPU to antother (and this shouldn't be proprietary!), but correct me if I'm wrong!

Does anyone know the bandwidth of cudaMemcpy on the current gen cards (GTX480 or 580)? All I know is that on the 5970 I'm able to transfer data from a GPU to another in AFR mode with a bandwidth of 4 GB/s. (ATI drivers do that internally)

Using the CPU/system memory is really suboptimal, and AFR too (Especially with 3-4 GPUS)!

[Edited by - OctavianTheFirst on January 6, 2011 12:05:21 PM]
DieterVW
DieterVW
The problem probably lies more in the way the graphics stack is setup on windows as well as some history behind it. Direct Compute is build on Direct3D so the terminology and design is inherently graphics based. Therefore there are some differences between DirectCompute and something like cuda which may not be expected if your background knowledge started in cuda.

SLI and CrossFire are proprietary solutions to making two graphics cards look like one. Windows does not naively support this for graphics -- it just wasn't designed into the graphics stack. So the drivers are doing all the work to share information across the PCI-E bus between the two or more adapters. The OS isn't involved and the DirectX API is completely ignorant of what's happening.

Traditionally I'm not aware of any default SLI or CrossFire setup for unknown graphics applications . The driver is actually coded with special paths to handle known software so that SLI or CrossFire can be used in a compatible manner. That's because it takes some profiling and testing to figure out how to leverage multi-gpu's without causing some sort of regression in visuals or performance.

Multi GPU's are also a thin slice of the market which means the windows graphics stack still hasn't been updated to support them natively and so the DX API contains nothing to manage such solutions. You'll have to hope for a proprietary API from ATI or NVIDA instead. I don't think that either company even has such extensions available to enable this for Directx. Instead those companies release their own solutions such as stream or cuda. Those technologies were designed with multi-gpu systems in mind. They are also proprietary.

Native OS support for such systems would probably look different. They wouldn't likely need the special connector linking the GPU's together. I don't know the details but it's possible that they use that connector as the BUS between cards instead of the PCI-E BUS. The OS would also have to be updated to do more intrusive resource management on the GPU's in order to figure out synchronization of resources. Plus, there would have to be some agreement on how adapters made by different manufactures would work together in these scenarios. Since right now it's all done by the driver I'd imagine that this wouldn't be simple task especially if the approaches have some fundamental differences.

CUDA doesn't use the graphics stack as far as I know. The adapters are seen as non-graphics parts. This key difference reduces a lot of restrictions required by the OS for graphics parts. NVIDIA has a paper on how they get better transfer speeds with CUDA because they can cut out all of the OS GPU management stack.

So I'm not arguing that it's impossible, I'm just that it has only been done independently by ATI and NVIDIA using their own hardware and software solutions and that there are a number of reasons that DirectX doesn't have explicit support for such functions or hardware configurations.
OctavianTheFirst
OctavianTheFirst
Thanks for all the info DieterVW!

Then for now, here's what I can do:

http://developer.download.nvidia.com/compute/cuda/sdk/website/samples.html#simpleD3D11Texture

And for ATI lovers all I found was OpenCL to DirectX10, but the 11 should follow:

http://developer.amd.com/support/KnowledgeBase/Lists/KnowledgeBase/DispForm.aspx?ID=85

I hope technology & oses will evolve to allow this kind of things as part of the API. As a coder I'd like to use the most of the processing power inside my pc, and i start thinking that, even for rendering, the best performance can be achieved through full control of each gpu. AFR adds lagg, and the others simply don't work. Let me code mine! Trying desperately to make 2-3-4 GPUs look like one (with the whole command queue & driver), and then adding the "multithread wanna be DeviceContext" doesn't really sound like the best solution. Sometimes less is more! Still, you explained well enough the reasons behind!

Finally, I'm really happy to see that DirectX evolved to include DirectCompute in the API. Hopefully at some point it will include a more general & programmable way of using multiple GPUs.

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.