Skip to main content
GameDev.net gamedev.net
🔒 Locked

Rare zlib crash (c++, linux)

Started by _Kami_ Mar 16, 2011 at 3:35 PM 5 replies 4.5k views
Original Post
_Kami_
_Kami_
Hi.

I have a very rare crash that happens sometimes when our servers are under heavy load. Its very hard to reproduce, and i have been unable to catch it in a debugger so far. (the debugger callstack you see here is automated and connected to a live service)


int net::IO_HandlerBase_t::AppendToSendQueue(const void *pData, uint nDataSize, bool bFlush)
{
if(m_pSendZStream != NULL)
{
m_pSendZStream->avail_in = nDataSize;
m_pSendZStream->next_in = (Bytef*)pData;
s_cStats.nSendPreCompress += m_pSendZStream->avail_in;

do
{
uint8 aCompressionBuffer[32*1024];
m_pSendZStream->avail_out = (uint)sizeof(aCompressionBuffer);
m_pSendZStream->next_out = (Bytef*)aCompressionBuffer;
int nResult = deflate(m_pSendZStream, bFlush?Z_PARTIAL_FLUSH:Z_NO_FLUSH); // line 550 (crash)
if(nResult!=Z_OK && nResult!=Z_BUF_ERROR)
{
return -1; // disconnect
}
s_cStats.nSendPostCompress += m_pSendZStream->next_out - aCompressionBuffer;
m_cSendQueue.PushBytes(aCompressionBuffer, m_pSendZStream->next_out-aCompressionBuffer);
}
while(m_pSendZStream->avail_in != 0);
}
else
{
m_cSendQueue.PushBytes(pData, nDataSize);
}

return 1;
}


relevant callstack:
#8 0x0000003062c0769e in _tr_flush_block () from /usr/lib64/libz.so.1
#9 0x0000003062c05335 in ?? () from /usr/lib64/libz.so.1
#10 0x0000003062c040d2 in deflate () from /usr/lib64/libz.so.1
#11 0x0000002a9669434c in net::IO_HandlerBase_t::AppendToSendQueue (this=0x2ab6d7b560, pData=0x0, nDataSize=0, bFlush=true) at IOHandlerBase.cpp:550
Now, i've read the zlib docs, and it seems our code does not handle avail_out == 0. As far as i understand from the docs, if avail_out == 0, then we should retry deflate() call, only with a bigger buffer. I would expect that if this was the cause, this would be some sort of infinite loop though.

Does anything obvious pop up looking at this code? like, omg, you are forgetting to do this&that.

Any help would be appriciated.
www.ageofconan.com
flodihn
flodihn

int net::IO_HandlerBase_t::AppendToSendQueue(const void *pData, uint nDataSize, bool bFlush)
{
if(m_pSendZStream == NULL) {
m_cSendQueue.PushBytes(pData, nDataSize);
return 1;
}

m_pSendZStream->avail_in = nDataSize;
m_pSendZStream->next_in = (Bytef*)pData;
s_cStats.nSendPreCompress += m_pSendZStream->avail_in;

while(m_pSendZStream->avail_in != 0) {
uint8 aCompressionBuffer[32*1024];
m_pSendZStream->avail_out = (uint)sizeof(aCompressionBuffer);
m_pSendZStream->next_out = (Bytef*)aCompressionBuffer;

int nResult = deflate(m_pSendZStream, bFlush?Z_PARTIAL_FLUSH:Z_NO_FLUSH); // line 550 (crash)

if(nResult != Z_OK && nResult != Z_BUF_ERROR)
return -1; // disconnect

s_cStats.nSendPostCompress += m_pSendZStream->next_out - aCompressionBuffer;
m_cSendQueue.PushBytes(aCompressionBuffer, m_pSendZStream->next_out-aCompressionBuffer);
}
return 1;
}


Is deflate your own function? To me it looks like the crash happens in there.
frob
frob
Since it only happens under heavy use and the problem is with stack-allocated memory, have you considered the possibilities of stack corruption or of stack exhaustion?
hplus0603
hplus0603

I have a very rare crash that happens sometimes when our servers are under heavy load. Its very hard to reproduce, and i have been unable to catch it in a debugger so far. (the debugger callstack you see here is automated and connected to a live service)



It's highly unlikely that the bug is in zlib itself. It's more likely that you are corrupting some data that it depends on. If you only see it during heavy load, then there's a few things that are more likely to be the culprit:

1) You're using threads or other asynchronous mechanisms (signals, etc), and there's a race somewhere.
2) You're using memory after it's freed, and under light load, it doesn't matter, because the memory isn't yet re-used.

Additionally, if it's stack memory that's corrupted, then you may have a stale pointer to the stack somewhere. I've seen this for example in this kind of use case:


class A {
public:
A(int val) : val_(val) {}
int val_;
};

class B {
A &foo_;
public:
B(A& foo) : foo_(foo) {}
};

B * function() {
A a(3);
return new B(a); /* dangling stack pointer hidden in the reference! */
}


The "B" lives for a long time, and when it reads/writes the "A" instance, it may smash a part of the stack that's currently not used, or it may smash a part of the stack that is used, thus being hard to trace.

Here's where C++ really allows you to shoot yourself in the foot, and doesn't have a lot of great tools to track the problems down. You could do things like scan each object returned for pointers to stack above the current stack frame, everywhere. You could use a debugging free, or a overridden operator new/delete, which nukes all memory to 0xcafefeed or similar, to flush out stale references. You could add random amounts of padding on the stack with alloca() and run a stress test to try to flush out the bug.

To debug this problem for real, I would seriously consider running a production server with something like VMWare replay debugging, and wait for the problem to happen. Then you have a record of all the events that went into the crash, and you can then look back in time to see how it came to be. Setting up replay debugging is annoying, and it may or may not work with your server OS, but it's probably worth a shot.

Bugs like these are why I find myself more productive in environments like Erlang, Node.js, C# or even PHP. (Although those systems have their own problems)
enum Bool { True, False, FileNotFound };
_Kami_
_Kami_
thanks for the replies.


Is deflate your own function? To me it looks like the crash happens in there.
[/quote]

Deflate is a zlib function api call. I find it hard to belive the actual bug is there since it has proven itself in many applications.


As mentioned by others, I doubt the bug is in zlib - but was hoping it was a api usage error (that would have been way easier to fix).

I will follow your advices about using memory tools to track down the problem. I have some thin stressclients, and will try to reproduce with server running in valgrind and hope I am able to catch the error runtime.
www.ageofconan.com
_Kami_
_Kami_
As suggested by others in this thread, this isse boiled down to beeing a race issue.

This code was beeing called by a different thread:

m_pSendZStream = (z_stream*)malloc(sizeof(*sendStream));
memset(m_pSendZStream, 0, sizeof(*m_pSendZStream));
sendStream->zalloc = zlib_alloc;
sendStream->zfree = zlib_free;
deflateInit(m_pSendZStream, 1);


As you can see from my first post and this, both are using the m_pSendZStream without semaphores, causing all kinds of problems.

strange that valgrind didnt catch it though.

Anyway, just putting the solution here for future ref, if someone manages to do the same thing and get a unexplainable crash in zlib;)

(this topic can be closed now)
www.ageofconan.com
Antheus
Antheus
I've used zlib for heavily threaded networking (albeit with asio, using thread-local buffers) but never noticed a problem with zlib itself.

I do seem to recall however, that zlib allocates a state buffer (~256kb) which must be maintained as thread local. I don't recall the details, but it might be that whatever zalloc points to must not be shared across threads, since it contains active (de)compressor state.

Semaphores/mutex/critical section and similar are highly undesirable. Zlib is very slow compared to rest of networking and will cause the server to degenerate to single-threaded.

From manual:
zalloc must return Z_NULL if there is not enough memory for the object. If zlib is used in a multi-threaded application, zalloc and zfree must be thread safe. [/quote]

This does not mean they must be within a semaphore, but that memory they return must not be used across threads - it contains running state of that particular stream's compressor. So for block compression, where entire buffer can be processed completely at once, using simple thread-local or even stack-allocated buffer is enough.

For streaming compression, memory footprint can become a problem, since it adds 64-256kb per connection.

"Thread-safe" is an ill defined word, allocator here needs to have transactional semantics with regard to each stream.


I also seem to recall that sequence of init calls matters. Creating and destroying zlib state is very expensive and should not be done for each step. The code I have does inflateInit to initialize the buffer once, but then uses inflateReset and inflateEnd each time compression needs to be done. This reuse complicates buffer handling a bit, since Init takes most of the time and it makes sense to reuse it.

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.