I've been debugging some code for a few days now. I don't get what the issue could be.
The code uses atomics and multiple worker threads to handle multiple jobs. It recently started crashing during thread cleanup (or maybe I just never triggered the bug before) on some runs (by that I mean the C++ standard library functions that get called after a thread terminates). The errors range from tcache_thread_shutdown(): unaligned tcache chunk detected to stuff like munmap_chunk(): invalid pointer. Sometimes, it is a straight up assertion that fails and spews itself into the console log. Anyway, I only use sequentially consistent atomics everywhere, and I test it in -O0 to rule out weird optimisation stuff or code reordering as the cause.
A short research seems to indicate that apparently when a thread shuts down, a heap sanity check is performed, which would then trigger on detected heap corruptions. So I went on to debug my memory using valgrind, and it could not detect any out of bounds accesses. Then I made sure there are no use-after-free allocation aliasing bugs by routing all calls to free() through a function that does not free the memory.
I also rerouted all calls to memory allocation through a function that adds half a kilobyte of padding to the front and back of each allocation, and uses magic words to fill the padding, and also a special first padding word that ensures I don't pass some modified pointer to free(). It also initialises all allocated memory to 0xbabababababababa to make sure this is not a case of uninitialised memory. Additionally, I modified my free() wrapper to check that the padding is intact and the first magic word value is found at the start of the padding, so that I can also rule out that the problem is caused by invalid pointers.
After adding the padding to all allocations, I was no longer able to trigger the memory corruption error messages on thread termination, but it could also just be a timing issue or something. And I still never got valgrind to complain about my memory usage, although that also isn't surprising.
What else could I do to debug this code? The multithreaded code is a very simple worker thread that takes a ring buffer of inputs and a ring buffer of outputs, and tracks the following variables via sequential atomics: number of issued jobs, number of processed jobs, whether to quit the worker loop. My ring buffer accesses are bounds-checked.
I'm thinking about replacing malloc() with a function that uses mprotect() and forcibly surrounds each allocation with a page of unaccessible memory, and either places all allocations at the start of the allocated area, or right at the end of it, so that any reads beyond the allocations should be impossible to miss. Maybe even using mmap() or something to use virtual addresses that are very far from each other.
Another weird thing is that rarely, the worker threads seem to get stuck or something while I poll for results, so that the polling does not terminate. However, since I only use sequential atomics, I don't get where the nondeterminism could be coming from. I also don't have work stealing or anything like that, and always use the same inputs for my program, and also the order in which jobs get passed to the workers is deterministic, as is the order in which I poll results from the workers. So, the only nondeterminism should be whether some thread maybe finishes one or two jobs between polls from the main thread, since the poll operation can fetch multiple results at once.
This is probably the most frustrating bug I ever had. I ran my program nonstop in GDB in a loop for thousands of times and could not get it to ever reproduce the heap corruption errors, so I also cannot debug them. I also tried compiling with -O1, which just changed the frequency at which the program gets stuck.
For context, here's the pipeline algorithm: Only the main thread issues new work, and there is only one worker thread per pipeline. The main thread is also the only one who polls results, so the results_consumed variable is not atomic, as the worker thread never accesses it. I separated all atomics into separate cachelines. I really don't think that this is the source of my problems. And I already did everything I could think of to fortify the rest of the code or to try observe the bug properly. I hope I didn't make a mistake transcribing and stripping down the following pseudocode:
struct Pipeline:
const Job[] job_queue
const Result[] result_queue
const u32 queue_cap
u32 results_consumed := 0
---- Cacheline boundary ----
atomic u32 jobs_issued := 0
---- Cacheline boundary ----
atomic u32 jobs_processed := 0
---- Cacheline boundary ----
atomic bool quit := true
worker_thread():
quit := SEQ_CST (false)
u32 _processed := (SEQ_CST jobs_processed)
while(! (SEQ_CST quit))
{
u32 _issued := (SEQ_CST jobs_issued)
if(_issued == _processed)
continue
while(_processed != _issued)
{
work(job_queue[_processed % queue_cap], &result_queue[_processed % queue_cap])
jobs_processed := SEQ_CST (++_processed)
}
}
quit := SEQ_CST (false)
issue(Job job):
u32 _issued := (SEQ_CST jobs_issued)
assert(_issued - results_consumed < queue_cap)
job_queue[_issued % queue_cap] := job
jobs_issued := SEQ_CST (_issued+1)
(Result[], u32) poll():
if((SEQ_CST jobs_issued) == (SEQ_CST jobs_processed))
return (null, 0)
while((SEQ_CST jobs_processed) == results_consumed) { ; }
u32 start := results_consumed % queue_cap
u32 end := (SEQ_CST jobs_processed)
if(end > start)
{
results_consumed += end-start
return (result_queue+start, end-start)
} else
{
results_consumed += (queue_cap-1)-start
return (result_queue+start, (queue_cap-1)-start)
}