Original Post
Ok fellow coders, lets see if we can come up with a fast way to alpha blend 2 32-bit pixels using Pentium assembly WITHOUT MMX. NO C allowed, assembly only.
The formula:
newpixel = (sourceRGB * alpha) + (destRGB * inverse alpha)
---------------------------------------------
256
Now, I''ve tried the brute force method of multipling each individual RGB component by the alpha value, dividing by 256 (with a shr,8) and then packing the components back up. This obviously sucks cause it would require 9 muls for the entire operation. Not gonna happen.
The second method I had in mind was doing a psuedo-packed integer register method. In this method we have a source ARGB quad sitting in EBX and an alpha value sitting in EAX and the idea is to perform a packed multiplication and control the overflow.
EBX - 00RRGGBB
EAX - AAAAAAAA
Then I mask out the RR and BB components in EBX using a mask of 0x00FF00FF and do the same to EAX. This results in this:
EBX - 00RR00BB
EAX - 00AA00AA
I then MUL EBX and the result in EAX is:
EAX - AARRAABB
which by the way means
(alpha*red) - HI EAX
(alpha*blue) - LO EAX
After a SHR EAX,8 and a MOV ECX,EAX we should end up with:
ECX - 00AR00AB
Next, we reload the original values back to EAX and EBX:
EBX - 00RRGGBB
EAX - AAAAAAAA
Only this time, I mask out the GG with a mask of 0x0000FF00 and do the same to EAX. Leaving:
EBX - 0000GG00
EAX - 0000AA00
I again MUL EBX and the result in EAX is:
EAX - 00AAGG00 (alpha * green)
After a SHR EAX,8:
EAX - 0000AG00
and a OR ECX,EAX we should end up with:
ECX - 00ARAGAB
NOW, we do the exact same thing to the destination RGB pixel and store the result to EBX. Now we should have this:
ECX - 00ARAGAB (sourceRGB * alpha)
EBX - 00IRIGIB (destRGB * inverse_alpha which is 255-alpha)
Now we need to add these two components togther while preventing overflow so we do this:
AND ECX,0x00FEFEFE
AND EBX,0x00FEFEFE
Which masks out the low bit of each component. We do an ADD EBX,ECX and an SHR EBX,8 and that should leave us with new pixel value to write back.
I tried it, and it kinda works but is still a little screwy SO...whoever has some good ideas and is willing to share THEN let the coding begin!
~-------------------------------------------------------~
Fred Di Sano
System Programmer - Artech Studios (Ottawa.CA)
~-------------------------------------------------------~