Original Post
Hi.
I have a very rare crash that happens sometimes when our servers are under heavy load. Its very hard to reproduce, and i have been unable to catch it in a debugger so far. (the debugger callstack you see here is automated and connected to a live service)
relevant callstack:
#8 0x0000003062c0769e in _tr_flush_block () from /usr/lib64/libz.so.1
#9 0x0000003062c05335 in ?? () from /usr/lib64/libz.so.1
#10 0x0000003062c040d2 in deflate () from /usr/lib64/libz.so.1
#11 0x0000002a9669434c in net::IO_HandlerBase_t::AppendToSendQueue (this=0x2ab6d7b560, pData=0x0, nDataSize=0, bFlush=true) at IOHandlerBase.cpp:550
Now, i've read the zlib docs, and it seems our code does not handle avail_out == 0. As far as i understand from the docs, if avail_out == 0, then we should retry deflate() call, only with a bigger buffer. I would expect that if this was the cause, this would be some sort of infinite loop though.
Does anything obvious pop up looking at this code? like, omg, you are forgetting to do this&that.
Any help would be appriciated.
I have a very rare crash that happens sometimes when our servers are under heavy load. Its very hard to reproduce, and i have been unable to catch it in a debugger so far. (the debugger callstack you see here is automated and connected to a live service)
int net::IO_HandlerBase_t::AppendToSendQueue(const void *pData, uint nDataSize, bool bFlush)
{
if(m_pSendZStream != NULL)
{
m_pSendZStream->avail_in = nDataSize;
m_pSendZStream->next_in = (Bytef*)pData;
s_cStats.nSendPreCompress += m_pSendZStream->avail_in;
do
{
uint8 aCompressionBuffer[32*1024];
m_pSendZStream->avail_out = (uint)sizeof(aCompressionBuffer);
m_pSendZStream->next_out = (Bytef*)aCompressionBuffer;
int nResult = deflate(m_pSendZStream, bFlush?Z_PARTIAL_FLUSH:Z_NO_FLUSH); // line 550 (crash)
if(nResult!=Z_OK && nResult!=Z_BUF_ERROR)
{
return -1; // disconnect
}
s_cStats.nSendPostCompress += m_pSendZStream->next_out - aCompressionBuffer;
m_cSendQueue.PushBytes(aCompressionBuffer, m_pSendZStream->next_out-aCompressionBuffer);
}
while(m_pSendZStream->avail_in != 0);
}
else
{
m_cSendQueue.PushBytes(pData, nDataSize);
}
return 1;
}
relevant callstack:
#8 0x0000003062c0769e in _tr_flush_block () from /usr/lib64/libz.so.1
#9 0x0000003062c05335 in ?? () from /usr/lib64/libz.so.1
#10 0x0000003062c040d2 in deflate () from /usr/lib64/libz.so.1
#11 0x0000002a9669434c in net::IO_HandlerBase_t::AppendToSendQueue (this=0x2ab6d7b560, pData=0x0, nDataSize=0, bFlush=true) at IOHandlerBase.cpp:550
Now, i've read the zlib docs, and it seems our code does not handle avail_out == 0. As far as i understand from the docs, if avail_out == 0, then we should retry deflate() call, only with a bigger buffer. I would expect that if this was the cause, this would be some sort of infinite loop though.
Does anything obvious pop up looking at this code? like, omg, you are forgetting to do this&that.
Any help would be appriciated.