Skip to main content
GameDev.net gamedev.net
🔒 Locked

Location of data in the class

Started by JoshuaWaring May 20, 2014 at 8:02 AM 17 replies 3.2k views
Original Post
JoshuaWaring
JoshuaWaring

when I define a class with only these variables


float _m11, _m12, _m13;
float _m21, _m22, _m23;
float _m31, _m32, _m33;

is it safe to assume that to access _m13 from the pointer of _m11 I can


(&_m11)[2]

and

to get to _m21


(&_m11)[3]

to my understanding this should be allowed, but I was wondering if it'll always be placed in RAM like this or some systems might shuffle them for a reason I don't know.

JoshuaWaring
JoshuaWaring

I was also curious if the same could be said for functions arguments,


Vector4(float _x, float _y, float _z, float _w);

Would all these be in the same order and sequential?

Thank you for any help, it would be much appreciated.

BitMaster
BitMaster
It's not safe because the compiler is allowed to add padding between the data members.
Crypter
Crypter

As mentioned above, it is not safe due to possible padding between members. You can however use a nameless struct here to achieve the same effect,


class matrix {
public:
 
   union {
        float m[9];
      struct {
         float _m11, _m12, _m13;
         float _m21, _m22, _m23;
         float _m31, _m32, _m33;
      };
   };
 
// now m[1] = _m11, m[2] = _m12 etc.
};

It is important to note that nameless structures are not standard although a few modern compilers support it.

TheComet
TheComet

I've seen this used before in a graphics rendering library.

The compiler does not reorder any members, and I assume there would be no padding, because 32-bit floats align correctly in 32-bit and 64-bit addressable memory, but it's not written in stone - the compiler is allowed to add padding if it wants to.

"I would try to find halo source code by bungie best fps engine ever created, u see why call of duty loses speed due to its detail." -- GettingNifty
JoshuaWaring
JoshuaWaring

Not what I wanted to hear, but thank you anyway :3

there's no way to ensure no padding is there?

guess it's back to my float _m[3][3] then

BitMaster
BitMaster

The compiler does not reorder any members, and I assume there would be no padding, because 32-bit floats align correctly in 32-bit and 64-bit addressable memory, but it's not written in stone - the compiler is allowed to add padding if it wants to.


I think assuming that is not unhealthy. I'm not in the habit of hand-rolling my matrix libraries but I would probably assume no padding and then
static_assert(sizeof(matrix) == 9 * sizeof(float), "This platform does not deal well with padding; must be handled explicitely");
Crypter
Crypter

Not what I wanted to hear, but thank you anyway :3

there's no way to ensure no padding is there?

guess it's back to my float _m[3][3] then

Technically you can enforce a specific alignment or none at all via compiler specific #pragma options. For example, in MSVC, #pragma pack directive or G++ attribute((packed)). TheComet is correct though, floats are very well standardized and guaranteed to be 32 bits so it is most probably safe to assume proper dword alignment without padding. The only guaranteed methods however would be compiler dependent; that is, resulting to nameless structures or #pragma.

You can also just overload operator[] and call it directly.

Alternatively just don't do it. You are not gaining anything from doing this after all. It can be done, of course, but at what benefit?

NightCreature83
NightCreature83

Not what I wanted to hear, but thank you anyway :3

there's no way to ensure no padding is there?

guess it's back to my float _m[3][3] then

Using these kind of trick, will work but it is also how you write unmaintainable code. And code readability is often more important than ease of writing your code, because in six months time you are going to be wondering why you wrote it that way.

The other way of doing this is


class Matrix44
{
private:
    union
    {
        float f[16];
        float mm[4][4];
        Vector4  v[4]; //Assuming vector4 is stored as 128 bytes contiguously
    } matrix;
}

You can now access matrix.f as an array, matrix.mm as a 2D array or as an array of vector4, if you add 16 float variables in the union you get 16 slot access options.

alvaro
alvaro
Washu posted a very neat implementation here: http://www.gamedev.net/topic/391237-criticize-my-custom-vector-class/page-5#entry3598509

struct Vector3 {
  float x, y, z;
  
  float &operator[](int index) {
    assert(index >= 0 && index < 3);
    return this->*members[index];
  }

  float operator[](int index) const {
    assert(index >= 0 && index < 3);
    return this->*members[index];
  }

  static float Vector3::* const members[3];
};

float Vector3::* const Vector3::members[3] = { &Vector3::x, &Vector3::y, &Vector3::z};
Bregma
Bregma




was also curious if the same could be said for functions arguments,

Vector4(float _x, float _y, float _z, float _w);

Would all these be in the same order and sequential?

Definitely not. Most early compilers for PDP and related architectures (VAX, PPC) always pushed arguments on the stack in reverse order. SPARCs and IBM mainframes would always pass 4 float args like in the example above in registers. I believe many currently-popular compilers will push the args on the stack in-order unless there are free registers, in which case one or more are passed in registers, depending on the optimization settings and link visibility. Of course, the compile could choose to inline the entire function, in which case there are no parameters.

So, no, never ever depend on the arguments to a function being in the same order and sequential in memory.

Stephen M. Webb
Professional Free Software Developer
Pink Horror
Pink Horror




to my understanding this should be allowed, but I was wondering if it'll always be placed in RAM like this or some systems might shuffle them for a reason I don't know.

I haven't seen the variables shuffled around, but I have worked on a gaming platform where this compiled into something that either crashed or gave garbage data back (I can't remember which unfortunately). The compiler is not obligated to actually compile that into how it appears that it should work. In the C++ standard, pointer arithmetic is only guaranteed to work within the bounds of whatever array the pointers are inside. Non-arrays are treated as arrays of length 1.

mark ds
mark ds

Nevermind - didn't read the question properly!

Kian
Kian

Not what I wanted to hear, but thank you anyway :3

there's no way to ensure no padding is there?

guess it's back to my float _m[3][3] then

Using these kind of trick, will work but it is also how you write unmaintainable code. And code readability is often more important than ease of writing your code, because in six months time you are going to be wondering why you wrote it that way.

The other way of doing this is


class Matrix44
{
private:
    union
    {
        float f[16];
        float mm[4][4];
        Vector4  v[4]; //Assuming vector4 is stored as 128 bytes contiguously
    } matrix;
}

You can now access matrix.f as an array, matrix.mm as a 2D array or as an array of vector4, if you add 16 float variables in the union you get 16 slot access options.

I'm not sure that's permitted. That is to say, you can do it, but you should only access the last member of the union you wrote to. This class looks meant to be used by writing to one member of the union and reading from another as convenient. That's undefined behavior. You'd need an enum type or something similar to keep track of what the last write operation was, and then different access methods instead of direct access to enforce the check, and then what are you even using a union for in the first place?.

Aardvajk
Aardvajk

Yep. It's undefined to read from a union other than the last member written to. Funny how everyone knows that compilers can insert padding in class members, but think that you can do this with a union with impunity. Believe GCC offers as a non-standard extension a guarantee that you can treat a union like this but it isn't standard C++.

All this stuff, one has to ask, is it really worth it to be able to say v.x as well as v[0]?

JoshuaWaring
JoshuaWaring

Just so people know, the reason the consistency is important isn't only nice use of variable names, but mainly SSE when I call _mm_loadu_ps(&_m11)

it'd need to be consistent to load in the m11, m12, m13 and m21 (which we'd then change back to 0 unless it's beneficial such as an addition) variables following

Hodgman
Hodgman

I'm not sure that's permitted. That is to say, you can do it, but you should only access the last member of the union you wrote to. This class looks meant to be used by writing to one member of the union and reading from another as convenient. That's undefined behavior. You'd need an enum type or something similar to keep track of what the last write operation was, and then different access methods instead of direct access to enforce the check, and then what are you even using a union for in the first place?.

The standard says that, yes.
But in practice, every game engine uses these hacks, and every C compiler supports them tongue.png

In fact, on many compilers, this is the preferred way of reinterpreting memory from one type to another -- e.g.


float data = 2.0f;
int bits = *(int*)&data//the standard says no! (but it will probably work on most compilers anyway...)
 
union {
  float f;
  int i;
} data;
data.f = 2.0f;
int bits = data.i;//the standard says no! but certain compilers recommend that you do this instead of the above!

As for the math classes, it might seem nice to have different ways of accessing matrix elements exposed -- e.g. mat[0][0], mat.m00, mat[0].x, etc, etc...
However, assuming your math classes are implemented using SSE, then these alternative accessors are hiding huge performance costs!

You can't mix and match SSE-code and float-based code... well, you can, but there's a big performance cost in doing so; constantly moving data from SSE registers to float registers and back is a huge waste of time.
If you let the user of your math library easily pull individual floats out of your structures, then the user will do such things -- and then the user's code will be slow, because it will be mixing up SSE-based and float-based operations.

If you're going after performance, it might be better to make it easier for the user to stick with SSE-based operations, and make float-based code appear ugly, which will discourage the user from writing it wink.png Thus I'd recommend not aliasing your SSE types with arrays of floats - also you should provide ways for the user to extract a single value into a vec4 variable, e.g. so GetX() == vec4(x,x,x,x)

JoshuaWaring
JoshuaWaring

So you'd recommend storing the variables inside a __m128 instead of a float array and remove the need for the loadu ?

and have the getVariable function which if I remember correctly something the compiler basically just directs you to the memory instead of actually performing a function call in a optimized version.

frob
frob

So you'd recommend storing the variables inside a __m128 instead of a float array and remove the need for the loadu ?
and have the getVariable function which if I remember correctly something the compiler basically just directs you to the memory instead of actually performing a function call in a optimized version.

Doesn't quite work that way, since the values may or may not be in memory. They might be inside one of the processors. In a worst case such operations require pushing values out from a register back out to memory and again between processors for cache coherence.

Decide how you are storing values. Use it consistently.

There are sometimes ways to speed things up if you know details about the processors, but they are non-portable. Your code might take advantage of specific layouts and locations, but then you discover a specific processor performs horribly under some circumstances.

For switching between an m128 and float[4], why would you do that? If you decide to use SIMD operations and keep thing in the assorted MMX or XXM or other registers, don't mix and match with floats since you'll spend so much time switching modes. ... unless you happen to have very specific knowledge about the processor and the compiler that will work. It doesn't usually work that way, don't do it.

As for the union hacks, they exist and happen to work, but try not to use them. Officially you should not use unions the way people described above. In practice, on most compilers and hardware it happens to work right now. That is no guarantee it will work on other systems, nor does it guarantee it will work in the future. But if you are fine with writing code that relies on the behavior, and your goal is to write something that happens to work today without crashing (as most games are) rather than writing something that is going to last, write code as is best for the project.

But remember: All it takes is a compiler update and the code (which is inherently broken and technically illegal) and the hack will stop happening to work, and suddenly begin to fail spectacularly. It is a bug in your code waiting to happen, except you happen to know for those specific cases that the bug won't happen on the current system with the current compiler.

There is standards-defined behavior, implementation-defined behavior, and undefined behavior. Know which group your hack falls within. If your hack falls within implementation defined behavior, the implementation may change at any time. Any updates to the compiler might modify that implementation-defined behavior. If your hack falls within undefined behavior that happens to work, heaven help you.

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.