Hello everyone!
I want to explain to you a problem that I had and how did I solve it to get some feedback and check if my architecture might have been different!
You can find the code I am describing at the repo, the bskysim folder.
context
My master thesis is a Discrete-Event simulation of Bluesky, informed with empirical data. The simulation propagates posts on repost and they get stored per user their timeline, implemented as two stacks (src/timelines.zig), one that the user consumes the posts from and another the posts get stored for the next user sessions.
The thing is that I tried to make this as performant as possible, and as I sm running this on a big server with 1TB of RAM, so I multithreaded the simulation to consume as much RAM as possible.
The simulation needs:
- the network topology (does not change while running)
- the configuration (does not change)
- the state: user policies and distributions, and the timelines as described up there!
So when I create the threads, I pass the pointer to the topology and the configuration (arena allocated) and each thread loads the state, which uses an SMPAllocator.
This section is essentiallu the main.zig
problem
The simulation runs up to 1 million users, so there are a lot of timelines. Those are inherently not upper bounded, so despite assuming a certain capacity, the timeline might need more memory, so the SMPAllocator is needed if the timeline needs to grow bigger thab the capacity.
The SMPAllocator uses a Slab of memory as a VMA for small (less than 64K) allocations, but when it needs more memory it mmaps a whole Slab for that specific object, and that adds a new VMA for that process.
Apparently, linux has a vm_max_alloc of arround 1 million VMA. If running the 100K topology with 12 workers, timelines keep betting bigger and bigger and the max VMA gets surpassed, and the process gets killed by plenty of RAM available on the machine.
I considered maybe that I was using the SMPAllocator wrong, but I generate one and all the threads share them, which is what the documentations says, as far as I understand.
solution
After tinkering to create an allocator to use a big mmap at the beginning of the simulation and share it by regions, I though that C might have an allocator that I can link, and I discovered the jemalloc allocator, which was exactly what I needed! By allocating different memory sized chunks it reduced the RAM ussge as well as avoided the max VMA problem.
question
Did I interpret wrongly the SMPAllocator/chosen a poorly design? Or its that my use case was not supported for this allocator? Should Zig consider new allocators in the std if yes?
Thank you for reading all of this ![]()