Skip to main content
GameDev.net gamedev.net
🔒 Locked

Communicating with Sound Card

Started by Phyrrus52 Sep 20, 2007 at 6:13 PM 22 replies 2.2k views
Original Post
Phyrrus52
Phyrrus52
Hi, I'm attempting to create a type of sound library in C++. I think I've got a pretty good idea on how to do this, but I need to learn how to communicate with and send bytes to the sound card (Sound Blaster, I suppose) to be played. Can anyone direct me to a tutorial on how to do this or give any tips? Thanks.
popsoftheyear
popsoftheyear
Why do you need to communicate directly with a specific sound card? Are you using DOS? If not, what OS are you using? Either way there is a ton of information for you out there, but you need to be more specific.

Cheers
Phyrrus52
Phyrrus52
Perhaps that was a little vague. Sorry, I'll explain:

I would like to manipulate sound data in real time. In order to do this, I need to be able to directly control the bytes of sound data and send them to the speaker.

In other words, if I have data:

A5 ED 4A ED B2 EE 2D EE (in hexadecimal)

I want to be able to run algorithms on that data and then directly play it by sending that data to the speakers.

I'm personally running Vista, but of course I'd expect this to be compatible with more than Vista.
TheAdmiral
TheAdmiral
Well, you could use DirectSound.

By creating a DirectSound buffer, locking it and dumping your raw data in, you can get your sound-card to play just about anything you like. I hope you're aware, though, that at 44kHz, eight bytes will last only around a millisecond.

Admiral
Ring3 Circus - Diary of a programmer, journal of a hacker.
Phyrrus52
Phyrrus52
Thanks for the help!

About 1/44 of a millisecond, I think! [smile] However, of course, sound data is much longer than 8 bytes usually, which leads to my concern:

I want to be able to control the timing of when certain bytes are played to within milliseconds - relative to other channels. (So, for instance, I don't care when channels one and two are playing, as long as it's within +- 50 milliseconds, but I do want channel 2 to be, say, 5 milliseconds delayed from channel 1 when both are playing the same sound data) This is in order to try to create a binaural sound library.

But, with, say a data rate of 60 kb/s (too low, but fine for now), we're talking about 60 bytes per millisecond! Is that too fast to consider fine-tuning the timing on the fly by creating a buffer? I mean, that's 60 bytes a millisecond flowing into the buffer! [wow]
TheAdmiral
TheAdmiral
Quote:
Original post by Phyrrus52
But, with, say a data rate of 60 kb/s (too low, but fine for now), we're talking about 60 bytes per millisecond! Is that too fast to consider fine-tuning the timing on the fly by creating a buffer? I mean, that's 60 bytes a millisecond flowing into the buffer! [wow]

Well, I can't imagine you'd be doing yourself any favours by attempting to lock and unlock the buffer every millisecond, for the sake of 60 bytes, but if you implement a write-cache, there should be no trouble. In system memory, store up as much audio data from the stream as is feasible and write these batches to the sound buffer (in a single contiguous copy operation) periodically - one lock/unlock per second would be fine. The data throughput of memory-writes and the motherboard buses is far in excess of 60kB/s (we're talking several GB/s) so I really wouldn't worry about that.

Today's computers are capable of dithering and pumping a real-time audio stream at 44.8kHz without breaking an idle sweat. Testament to this is the fact that the data can further be FFTed, processed in a dozen different ways and IFFTed prior to dithering without pushing the CPU out of its comfort zone.

Admiral
Ring3 Circus - Diary of a programmer, journal of a hacker.
Gumgo
Gumgo
(I'm also working on this with Phyrrus52.)

We plan to store sound data in an array of bytes. The way we will achieve the binaural effect (as well as the doppler effect) is to constantly stream the bytes to the sound card except rather than just sending something like
SoundData[time]
we would send
SoundData[time-(distance/SpeedOfSound)]
Of course, we would send this two times: one for each channel and the distance is measured from the source of the sound to each "ear". Also, the distance/SpeedOfSound will be interpolated from previous position to current position.

Anyways, assuming that the time of each game cycle is 20ms, at 60kbps that would mean sending 3,000 bytes to the sound card every 20ms, per channel, per sound playing. So if 10 sound were playing using both channels, that would be sending nearly 3MB to the sound card each second (if I understand correctly and am not missing some other way to do this).

Anyways, I'll edit this when I get home to actually ask something...
Sneftel
Sneftel
Quote:
Original post by Gumgo
Anyways, assuming that the time of each game cycle is 20ms, at 60kbps that would mean sending 3,000 bytes to the sound card every 20ms, per channel, per sound playing. So if 10 sound were playing using both channels, that would be sending nearly 3MB to the sound card each second (if I understand correctly and am not missing some other way to do this).

Sound cards use DMA, whereby you tell the sound card where to find the data and it handles the copying itself without taking up valuable CPU time. Moreover, 3 MB/sec is not a significant strain on the bus; a modern PCIe bus can carry 250 MB/sec on each lane, and the memory itself can transfer several gigabytes per second.
Gumgo
Gumgo
Quote:
Sound cards use DMA, whereby you tell the sound card where to find the data and it handles the copying itself without taking up valuable CPU time. Moreover, 3 MB/sec is not a significant strain on the bus; a modern PCIe bus can carry 250 MB/sec on each lane, and the memory itself can transfer several gigabytes per second.

Oh good, then no problems there.

Only one more thing:
Quote:
Well, I can't imagine you'd be doing yourself any favours by attempting to lock and unlock the buffer every millisecond, for the sake of 60 bytes, but if you implement a write-cache, there should be no trouble. In system memory, store up as much audio data from the stream as is feasible and write these batches to the sound buffer (in a single contiguous copy operation) periodically - one lock/unlock per second would be fine.

Hmm, maybe I'm not understanding this but doesn't that mean that the sound would be delayed by one second? We were planning on sending data to the sound card every 20ms (the length of the game loop).
FritoBandito
FritoBandito
Check this out: (Google rocks!)

http://www.ddj.com/184410687

it's an old article, but worth a look.
Phyrrus52
Phyrrus52
Hey, thanks!

That's an interesting article, but it mostly has to do, it seems, with audio recording. An interesting topic though... I really should learn more about it!
TheAdmiral
TheAdmiral
Quote:
Original post by Gumgo
Hmm, maybe I'm not understanding this but doesn't that mean that the sound would be delayed by one second? We were planning on sending data to the sound card every 20ms (the length of the game loop).

You're understanding fine, I just meant that you should submit the sound data in advance of its use. In this case, you'd need to start filling the buffer at least a second before it is to be submitted.

All sound cards have a degree of latency: the minimum time taken for submitted data to reach the speakers. While very expensive sound cards have latency that isn't noticeable, most mid-range PCs will exhibit more than enough latency to dash any hopes you may have of submitting arbitrary sound data in real-time. This is the reason that all sound frameworks (e.g. DirectSound, FMOD) require you to create sound buffers before you plan to use them: this way, the card is all ready to dither the audio data at the drop of a hat.
Dynamic sound buffers (as I described at first) are the closest you'll get to real-time submission. Now I don't know the exact figures (which will be hardware-dependent) but I can't imagine it will be possible to submit data within say 100ms of it being playable. If your game pivots on this principle, then I'm afraid you'll have to redesign.

May I ask why you are reinventing these wheels? DirectSound offers a rich selection of 3D sound features, including binaurality and Doppler. There are also a few extras that you couldn't emulate with your proposed setup, such as realistic attenuation and reverb. Even better, the whole thing is implemented in a well-thought-out object-oriented fashion, meaning you have to do little more than create a 3D buffer and listener to get things going.

Admiral
Ring3 Circus - Diary of a programmer, journal of a hacker.
Phyrrus52
Phyrrus52
Quote:
May I ask why you are reinventing these wheels? DirectSound offers a rich selection of 3D sound features... ...including binaurality...


Whoa!

We're reinventing the wheel because... we didn't know that wheel existed! Never found any evidence of it in documentation. I don't suppose you could point me to it? That would be great!
RSC_x
RSC_x
hey please stop this thing...! [dead]
how you can reinwent an existing thing.?
you cant reinwent anything including whell.
this is just copying or learning.
when you open a door you are not going to invent the hall.
hall isa allready invented.
so you just learnin how to find it again.
let me quess next time you are going to invent dir command on dos?[embarrass]
© Loading...<font col
Gumgo
Gumgo
If you're not going to be helpful please don't post.
TheAdmiral
TheAdmiral
Quote:
Original post by Phyrrus52
We're reinventing the wheel because... we didn't know that wheel existed! Never found any evidence of it in documentation. I don't suppose you could point me to it? That would be great!

Perhaps I misled you a little: The binaural effect isn't directly supported, but it can be emulated about as effectively as is possible in real-time by duplicating buffers. All the other effects I mentioned do feature in the documentation.

A first-order binaural effect could be produced by creating each buffer twice, positioning them in the same place, but panning them to opposite sides. Now by triggering the two buffers to play at slightly different times (according to their position relative to the listener), the buffer will reach the two speakers separately. I'm not sure how much success you would have, but by adding EQ, delay and/or reverb to the two buffers, you could simulate the subsequent echoes and non-uniform attenuation required to represent the vertical dimension.

Now I don't know anything about your project, but unless your primary aim is to create a real-time binaural simulation, I fear you may be wasting your time. Of course, your target-audience is limited to headphones users - the natural acoustics of a typical room render the subtle effect inaudible under stereo speakers.

You may have far more success by approximating the effect using multiple recordings. If you could create several stereo buffers containing the same sample, but binaurally biased to originate from different directions (four compass directions at a minimum; up and down would be nice for 3D) then you could feasibly create the effect of sound coming from any direction by playing the samples simultaneously at the appropriate volumes, trigonometrically determined. This is far more scalable than the previous ideas, but it's a little memory-intensive.

Let us know what your goals and resources are, and perhaps we can help you decide which approach is best-suited.

Admiral
Ring3 Circus - Diary of a programmer, journal of a hacker.
Gumgo
Gumgo
Quote:
Now I don't know anything about your project, but unless your primary aim is to create a real-time binaural simulation, I fear you may be wasting your time. Of course, your target-audience is limited to headphones users - the natural acoustics of a typical room render the subtle effect inaudible under stereo speakers.

Yep, that's our goal.
Quote:
There are also a few extras that you couldn't emulate with your proposed setup, such as realistic attenuation and reverb.

Could we maybe use these alongside our (attempt of a) binaural setup?

As for real time... there's one thing I'm a bit confused about (maybe I'm thinking about this entirely wrong or something). If it is possible to change things like reverb, volume, flange, etc. in real time, what would be different about this? I mean, once the data gets to the sound card, that's what it plays, right? So why can other effects be changed in real time easier than this?
FritoBandito
FritoBandito
The amount of time it would take to resample the audio input and apply a reverb algo to it would preclude it from being realtime anymore.

You are already up against the rate at which a soundcard can generate an interrupt and get serviced, and this, i fear, would be simply too much.

TheAdmiral
TheAdmiral
Quote:
Original post by Gumgo
If it is possible to change things like reverb, volume, flange, etc. in real time, what would be different about this?

I'm not an expert on this, but my understanding is that effect plugins are written according to very strict standards so that they may be allowed to run lower down the food chain. That is, the processing is performed by the driver or (in the accelerated case) the card itself using DMA, interrupts and a whole host of other scary ring0 technology.

Admiral
Ring3 Circus - Diary of a programmer, journal of a hacker.
Kylotan
Kylotan
No, most effect plugins are standard user-level programs that just take a stream of data in and a stream of data out. You can change pretty much all these effects in 'real-time' providing you're not expecting zero latency (which you won't get on Windows anyway).

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.