Skip to main content
GameDev.net gamedev.net
✓ Solved 🔒 Locked

Are there commonly acknowledged guidelines on samplerates, bitdepth, file types and tracklengths for game audio?

Started by neonblackclouds Apr 28 at 5:56 PM 8 replies 1.8k views
Original Post
neonblackclouds
neonblackclouds

Hello everyone,

i have a few questions regarding common practices when it comes to game audio, especially when provided as a sample-pack or sound library. I am an audio person and do not have much background in game development even though I have a rough overview of the field and generally like video games.

Currently I am doing some sound design experiments and though it would be fun and maybe helpful to others, if I release some of it as a sound effects / background tracks pack. I listen to a lot of video game soundtracks and also did some research on some asset stores, but I am still unsure how to best deliver the files such that it is as easy/convenient as possible for game devs (who might not be super deep into audio engineering) to use the files.

So here are my questions (any additional advice is welcome as well):

1.) For background "ambience" type tracks (rain, drones, etc.) how long would they optimally be? I know that they should loop seamless, but when I checked some asset packs the tracks seemed all rather short to me (between 0:30 and 2:00). For background ambience I would have guessed the longer the better since more recognizable patterns would not stand out as much as in a 60sec track repeated x times. Are they that short to reduce the file size? A normal ambient track as a musical release could be well beyond 10min and I do not really understand why it should be shorter for games given that the filesize would not be crazy big for modern hardware? When checking soundtracks e.g. for the Amnesia games, the tracks are much longer than 2min. Is it misleading to look at them for comparison since they are probably rearranged a bit for the OST?

2.) For sound FX: Obviously it would be a good idea to package the sounds in a dry (no additional e.g. room effects) version. Is a wet version (e.g. with some room reverb) also of use or is this something people will handle by themselves inside the engine? I thought maybe a dry version and 2 wet versions, one subtle and one a bit more out there would be fun, but I am unsure if it is useful to anyone.

3.) Is there a preferred samplerate, bitdepth and file type for such an audio pack? I can deliver really high quality uncompressed files (e.g. 96kHz, 24bit, .wav) for people to downsample it themselves later if necessary. Are there commonly used settings for game audio (like the 48kHz in film) or does it not matter because people compress it later to reduce filesizes anyway (read something like this in another forum post but was unsure how common a practice it is). Do people not really care as long as it takes up less space and would prefer a compressed version made from e.g. a 44.1kHz (CD quality) source to be included for direct use? The asset packs I looked at mostly had uncompressed or lossless compressed files but the samplerates where a little bit all over the place. Granted most of them were for free so maybe the quality standard was not very high for some of them.

4.) How loud are things expected to be? Do I master the background tracks like I would for e.g. streaming (-9 to -14LUFS)? Are sound effect samples expected to be normalized?

5.) This one is a bit off-topic but what asset stores / pages would you recommend putting stuff like that on? I want to release it for free, so I am not willing to pay any fees (or at least no big amount of money).

I know this is a lot of text and a lot of questions. If anyone would like to give their opinion on just one or two questions that already is very helpful. Thank you in advance for your time.

Accepted Answer
  1. I aim for minimum 30 seconds to a maximum of 2 minutes for my ambiences (environmental ambient sounds like wind/rain, not music). Any shorter and you can notice looping, any longer and you are wasting space. If I have more than 2 minutes of audio, I will split it up into multiple clips that are randomly selected (usually not more than 2 separate clips). For ambient music however, I have some tracks over 10 minutes. Longer is better there, but you may also want lots of shorter clips to splice together to make it more dynamic and interesting.

    For selling assets, you should probably provide both the original source files of maximum possible length, plus a few curated loops of shorter lengths that are cut from the master file. This allows each developer to select custom loops if they want, but also provides something that will work out of the box.


  2. No reverb at all in any sound effects. Reverb should always be applied in the engine, because the ratio of dry to wet is an important psychoacoustic cue for distance to the source. Reverb is usually roughly constant within a room, but the direct sound varies with 1/distance from the source, meaning the direct:reverberant ratio is higher closer to the source. For this reason when recording audio I will either record outdoors (near anechoic conditions), or in a well-treated dead studio room, never in a reverberant room. Getting the microphone as close as possible to the source helps get better SNR.


    You didn't ask, but in-game point audio sources should always be mono. Avoid "fake" stereo sounds or 2-channel binaural sounds unless they will be played head-locked. "Stereo" should come from the game's 3D audio spatialization system applied to mono sounds at different locations. For ambiences, the state of the art is ambisonics. This format captures full-sphere audio with 4 or more channels, supports head rotation, and can be decoded to binaural or any surround format in-game. You will need an ambisonic microphone to record those. I use a Core Sound OctoMic.


  3. Most engines will have their own audio encoding pipelines, which theoretically include sample rate conversion (I don't know how good the quality is). I record and author my sound effects at 96kHz, 32-bit float .wav. I would recommend providing the highest quality format that you have. In my engine in the final game build the audio is converted to:

    * For mono sound effects and stereo music, 44.1 kHz Ogg Vorbis.
    * For ambiences in ambisonic format, 48kHz Opus (because it only supports 48kHz, and is the only freely available ambisonic compression format).

    Keep in mind that if your audio doesn't match the sample rate of the audio engine, then those audio clips will have to be resampled on-the-fly which uses CPU (or harms quality if a bad algorithm is used). It is ideal to run the audio engine at a fixed rate (I use 44.1kHz), and convert sound effects to that rate when the game is built, to avoid runtime sample rate conversion. The only conversion should be at the final device output, if it doesn't match the game's rate. If you use any runtime pitch shifting, that will also cause sample rate conversion.


  4. For music and interface sounds, I would master those to be at a comfortable level with minimal master compression/limiting (like you would deliver to a mastering engineer). In a game, the mastering chain is applied at the very end, after mixing in the game audio, so all audio should be "pre master". The loudness doesn't really matter for in-game sounds because there are multiple layers of gain control and attenuation curves. Furthermore, there is no quality loss from low signal levels or clipping from high signal levels in common formats like Ogg Vorbis, they are inherently floating point, so level doesn't matter at all, unlike with integer .wav files. I would stick to 32-bit float .wav to avoid concerns about clipping/low levels.

    For my game, it is very important to keep the relative levels of sound effects consistent, so that I get as realistic a sound pressure level in the game as possible. For this reason I do NOT normalize sound effects, and I take special care to ensure that all effects I record of a specific type have the correct relative loudnesses. For example, for footsteps I keep all footstep sounds at the same gain, regardless of whether they are loud or soft footsteps. Then, I can be sure that the loud and soft footsteps are played at the correct relative levels in game, without any extra gain. If I normalized them, I would have to undo the normalization gains later to correct that. However, I do apply some rough normalization for different types of sounds (e.g. gunshot vs. footstep), so that they all peak near 0 dBFS, regardless of real-world loudness.

    For each type of sound effect (e.g. footstep on gravel, gunshot) I also assign an approximate "source power" value in dB SWL (sound power level). For example, all footsteps are calibrated to play at 65 dB SWL, meaning that if an audio sample is normalized to +/- 1 (0 dBFS), it would have a sound power of 65 dB in the audio simulation. Gunshots on the other hand have a realistic power of ~130 dB SWL, 65 dB higher than footsteps. I use the power value to apply an extra gain to the audio source when it is played. This produces correct relative levels between different types of sources, and produces "high dynamic range" audio. Gunshots will really be 60dB+ louder than most sounds. For this to work in practice, you need a good mastering chain that can compress/limit very loud sounds down to 0 dBFS without serious distortion or pumping. I don't think most people are doing it this way, but I am trying to push the boundaries of audio realism in games (my main expertise).

frob
frob

It will depend on the project and their expectations.

Some are as you described, looking for CD quality, 44.1kHz stereo. Some are expecting Dolby Atmos 7.1.4 192 kHz lossless masters.

Masters on professional projects are expected to be mixed and adjusted, so they should use the full audio range. On an amateur project they may do mixing or they may just be looking for mp3s to play directly, so that's going to vary.

Effects on the audio are going to be the same. In professional project they're more likely to mix and adjust them, or they may ask you to mix them, but either way the default would be to start dry and to discuss about how they'd like them. On amateur projects, more likely than not they're going to use them as-is.

Duration of tracks is going to depend on the project, some will want long, some short, some varied.

So all across the board: it depends.

neonblackclouds
neonblackclouds

First of all: Thank you to the both of you for taking some time out of your day to answer my questions so thorougly.

@frob: So my takeaway here would be that it definitely does not hurt to give people options? Since I will just put this out for people to use, I assume there will not be much of a back and forth regarding the audio needs of specific persons. I just wanted to make sure that I do not miss anything like a industry standard that would render a lot of my (down-)conversion and "adding effects" work completely useless. Since this would be my first compilation of material I guess that it also is aimed more towards amateur projects, so giving preprocessed options might even be more useful in that case.

@Aressera: Thank you so much for taking your time to write such a detailed answer. It is highly appreciated. Thank you for pointing out that much of the audio material is expected to be mono for point audio sources. Makes sense. I checked out wwise and fmod a bit so I see how this would be used to place them in a panorama that is dynamic during runtime. I will definitely take care of that (most of the effect sounds from recorded material are mostly mono anyway since I predominantly only use one mic for them). I checked out ambisonics (previously I was unaware it exists). Really cool technology but I do not know if I am ready to go down that rabbit hole any time soon. I also use 96kHz 32bit float .wav as my source format so I will make sure to add this as a high quality, high-disk space version. Regarding 4.) the way to go would then be to process similar sounds as a group to keep their internal dynamic range intact? Definitely makes sense from an ease of use perspective. Thank you for pointing that out. Regarding the last paragraph: While I see the appeal this probably is far from "common practice" correct? This is beyond the scope of just preparing audio for this process anyway but definitely something I never really thought about. It seems to me that mostly games go for convenient rather than realistic dynamic range such that somebody talking could still be heard over e.g. an explosion/gunshot. Anyway thank you for giving me all this information which will I will probably need some more time to fully digest. Very insightful.

Another related question that just came to my mind: What naming convention do you tend to stick to for delivered audio? Do any of you use something like UCS and is it currently even widely enough adopted to make sense?

frob
frob

neonblackclouds wrote:

I assume there will not be much of a back and forth regarding the audio needs of specific persons.

Again, who you are working with varies tremendously.

And what you bring to the project also matters tremendously.

If you're posting to SoundCloud or BandCamp or something, you're putting out the final music. You're also going to be invisible due to everybody who ever downloaded a sequencer thinks they're a good composer.

If you're working directly with a studio, you're giving building blocks and everything is custom to their specs. You'll be working with a producer and likely an audio director. But even then, that's not something that is likely. Who you are, what you've done, and your background matter.



Now for the hash truth. From the description, I don't think you're experienced enough for the studio work.

If you're working with a studio instead of just a "buy whatever is published", the expectations are for the composer to be a top-notch performer, everybody I know is not just a good composer but stellar performance pianist which they use during their composition.

Further, who you know, and networking matter, in many ways they matter more than anything. If I personally want a composer, I'm picking an experienced composer I know who I've worked with before, who I happen to know has a degree is music production, over a decade of real-world experience, and is an amazing live performer at clubs, events, and such. He knows his worth as a composer, with an extensive portfolio and IMDB page. If my first choice wasn't available, I've got a second and third choice, both skilled and trained in music performance rather than music production, and after those three I'd likely turn to other veterans like Nate Madsen (a moderator on the site) with university degrees in music performance and decades of game music experience. My first choice has "reasonable rates" that most indies could never afford, yet reasonable for the amazing quality I know I'll get, but also no point approaching him if it's not a 5-figure deal, 6-figure preferred but I know we'll be working together for several months. My second choice is downright cheap at 70/hr for the work he does, might be a 'friends' rate as a contractor, it will average out to a few hundred bucks per finished track.

For each one of those I've worked with, I know we can show what we want it to go with and they can immediately play out themes on the keyboard. Interrupt them with "do that again in a jazzy style", "maybe a techno stye", or "again but in rock", "good, but heavier", or ask "can you make it a bit more driven" and they'll be beatboxing a drumline right along with his improvised keyboard playing to themes, then a bit later after we've iterated a bit he'll have the thing worked out as a complete symphonic mix, with what to him is a rough first pass, but is still better than most of the "music packs" people post about and share. It isn't just "I have a music editor and can make songs", it's first and foremost "I'm an amazing composer and performer", which makes all the difference in the finished music.

For audio effects, we'll be sharing the project in wwise and simply hiring them for a few months of contract work. Saying that you've dabbled in wwise and fmod is nice and shows you've got a little interest, but it would also exclude you from most of the work, just too inexperienced. We'd need someone we could link to Perforce, point them to the wwise project to and they could fill everything out, get the system running within a day. That isn't someone who dabbles in it, that's someone who lives it every day as their career.

For professional audio work I'm not turning to an unknown random person after listening to a dozen hours of music. I'm turning to my network of experienced composers and performers.

You might still be able to sell music by posting it as an audio pack or something, but it's unlikely. There are already millions of hours of royalty-free sounds out there, and even AI sites like Suno or OpenMusic offer full commercial rights to generated songs -- and they're relatively good compositions -- for just a few bucks. For most people who need music, they'll either need it as part of a professional studio or they'll turn to other sources, they won't be paying for hobby musicians.

neonblackclouds
neonblackclouds

Hi frob,

In my initial question I wrote:

neonblackclouds wrote:

i have a few questions regarding common practices when it comes to game audio, especially when provided as a sample-pack or sound library. I am an audio person and do not have much background in game development even though I have a rough overview of the field and generally like video games.

So I think it goes without saying that I am not experienced enough for studio work, at least not in the capacity you are describing. This is not a "harsh truth" I needed to hear, neither did I write that I am looking to be hired. In fact, I wrote that I want to put something out for free, because I am experimenting with sound design. I am not even much of a composer and never said so.
As a person rather new to the field of audio for interactive media, I was asking what I thought of as rather reasonable questions about technical standards and preferences.

I do not understand where most of your second answer comes from. Even if all of it might be true and is great for you and the talented people you can pay properly, parts of it reads to me as rather condescending and above all unprompted. What am I supposed to take away from this?

In any case, I think I got some valuable technical insight and for now will proceed with my project.
Thank you for your time everyone 🙂

frob
frob

We get people every few days posting their wares as a new composer, new sound effects packs, and similar. Most are objectively terrible, with the people having no understanding of music, their packages effectively adding clutter and noise rather than meaningful content to the pool.

It is a frustration, amateur work is important as nobody starts as an expert, but the final bullet point in the questions is triggering. They are great to share with family and friends, but not really on the commercial sites.

neonblackclouds
neonblackclouds

I see how that can be frustrating and I also see all the locked self-advertisement posts in the audio forum that make it hard to use, but this is explicitely not what I was here for. I did not "post my wares" and I read the forum guidelines. It was an honest (and partly successful) attempt at getting some additional insight into a new topic.

If the final bullet point of my question was triggering you I am sorry, but a short explanation on why would have been much more productive. I also disagree on the friends and family thing. I think getting your stuff out there is part of any creative process and should not require the approval of anyone. Also from a practical point, the feedback from friends and family tends to be less objective and/or competent.

Look, I can not speak for the general experience in your field of work, but I have been an audio engineer in small clubs for many years and have seen so many absolutely unbearable bands and solo artists perform. (I have often heard engineers they occasionally brought along do really bad work too.) Still I would rather have them have a place to play their music than not, as long as they behaved themselves. Its how people improve.

On the other hand I understand that "professionals" need a space to communicate among themselves and I definitely did not want to annoy anyone (be it here or at "commercial sites" for assets). From what I saw in other forum posts on this page I thought this would be a place that also is for "newcomers" or however you want to call it. I would have never written a question otherwise. If this is a forum for "professionals working in the gaming industry" only, there honestly should be a disclaimer somewhere in the forum guidelines or along with the registration process.

I think I had enough arguing over the internet if thats ok for you. I hope you at least partially can see my point. I see yours but still felt the way it was expressed was not great and I also think your understandable frustration does not really have all that much to do with my question in particular.

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.