Skip to main content
GameDev.net gamedev.net
🔒 Locked

FAST software alpha blending

Started by fredo Jan 28, 2000 at 10:46 AM 2 replies 2.5k views
Original Post
fredo
fredo
Ok fellow coders, lets see if we can come up with a fast way to alpha blend 2 32-bit pixels using Pentium assembly WITHOUT MMX. NO C allowed, assembly only. The formula: newpixel = (sourceRGB * alpha) + (destRGB * inverse alpha) --------------------------------------------- 256 Now, I''ve tried the brute force method of multipling each individual RGB component by the alpha value, dividing by 256 (with a shr,8) and then packing the components back up. This obviously sucks cause it would require 9 muls for the entire operation. Not gonna happen. The second method I had in mind was doing a psuedo-packed integer register method. In this method we have a source ARGB quad sitting in EBX and an alpha value sitting in EAX and the idea is to perform a packed multiplication and control the overflow. EBX - 00RRGGBB EAX - AAAAAAAA Then I mask out the RR and BB components in EBX using a mask of 0x00FF00FF and do the same to EAX. This results in this: EBX - 00RR00BB EAX - 00AA00AA I then MUL EBX and the result in EAX is: EAX - AARRAABB which by the way means (alpha*red) - HI EAX (alpha*blue) - LO EAX After a SHR EAX,8 and a MOV ECX,EAX we should end up with: ECX - 00AR00AB Next, we reload the original values back to EAX and EBX: EBX - 00RRGGBB EAX - AAAAAAAA Only this time, I mask out the GG with a mask of 0x0000FF00 and do the same to EAX. Leaving: EBX - 0000GG00 EAX - 0000AA00 I again MUL EBX and the result in EAX is: EAX - 00AAGG00 (alpha * green) After a SHR EAX,8: EAX - 0000AG00 and a OR ECX,EAX we should end up with: ECX - 00ARAGAB NOW, we do the exact same thing to the destination RGB pixel and store the result to EBX. Now we should have this: ECX - 00ARAGAB (sourceRGB * alpha) EBX - 00IRIGIB (destRGB * inverse_alpha which is 255-alpha) Now we need to add these two components togther while preventing overflow so we do this: AND ECX,0x00FEFEFE AND EBX,0x00FEFEFE Which masks out the low bit of each component. We do an ADD EBX,ECX and an SHR EBX,8 and that should leave us with new pixel value to write back. I tried it, and it kinda works but is still a little screwy SO...whoever has some good ideas and is willing to share THEN let the coding begin! ~-------------------------------------------------------~ Fred Di Sano System Programmer - Artech Studios (Ottawa.CA) ~-------------------------------------------------------~
~-------------------------------------------------------~ Fred Di Sano System Programmer - Artech Studios (Ottawa.CA)~-------------------------------------------------------~
FlyFire
FlyFire
ok, let''s begin then

First of all, let''s review the formula:
res = src*alpha + dst*(1-alpha);
res = src*alpha + dst - dst*alpha;
res = (src-dst)*alpha + dst;

thus we have only one multiply instead of two.

Now about implementation (first one is simular to yours):

we mask out green value in both src and dst, substract dst from src and multiply it by alpha. Then add dst and mask out green field again.
The same for r&b.
After this, we can combine color back.

(Note, that in 15/16 bpp modes we can do alpha blending with only one multiply. example code:


x = ((x&0xFFFF)/(x<<16))&0x7E0F81F;
x = ((x&0xFFFF)/(x<<16))&0x7E0F81F;
res = ((x-y)*alpha/32+y)&0x7E0F81F;
0<=alpha<=31

(you can convert it to asm, i hope )

That''s how alpha blending can be done using multiplyes. But integer multiply is too slow. I''ve had an idea to use fpu for alpha blending, but i haven''t implemented it yet.

Another idea is to write specific code for each alpha value and do alpha blending using only shifts and adds. The only overhead is that you need a lot of code space (a few kb of code). But whole thing gives a great speed up!
But when you''ll start coding each code part, you''ll find some parts similar (the only difference is shift values), so they can be used for same thing). I haven''t experemented with this too much, but you can try to...



FlyFire/CodeX
http://codexorg.webjump.com

fredo
fredo
Umm, I think you may have made a mistake in your calculation. As far as 32-bit pixels go anyways.

You state that the alpha blend formula is this

res = src*alpha + dst*(1-alpha)

Is it me or is not supposed to be:

res = ((src*alpha) + dst*(1-alpha)) / 256

which would reduce to:

res = (((src-dst)*alpha) / 256) + dst

?

What do you think?

~-------------------------------------------------------~ Fred Di Sano System Programmer - Artech Studios (Ottawa.CA)~-------------------------------------------------------~
Starfall
Starfall
Actually, it depends on whether ''alpha'' is a value from 0 to 1, or a value from 0 to 255. If it''s 0 to 255 then I believe it should be

res = ((src*alpha) + dst*(255-alpha)) / 256

Otherwise the initial formula is correct.

Regards

Starfall

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.