Too many VMA with SMPAllocator

Hello everyone!

I want to explain to you a problem that I had and how did I solve it to get some feedback and check if my architecture might have been different!

You can find the code I am describing at the repo, the bskysim folder.

context
My master thesis is a Discrete-Event simulation of Bluesky, informed with empirical data. The simulation propagates posts on repost and they get stored per user their timeline, implemented as two stacks (src/timelines.zig), one that the user consumes the posts from and another the posts get stored for the next user sessions.

The thing is that I tried to make this as performant as possible, and as I sm running this on a big server with 1TB of RAM, so I multithreaded the simulation to consume as much RAM as possible.

The simulation needs:

  • the network topology (does not change while running)
  • the configuration (does not change)
  • the state: user policies and distributions, and the timelines as described up there!

So when I create the threads, I pass the pointer to the topology and the configuration (arena allocated) and each thread loads the state, which uses an SMPAllocator.

This section is essentiallu the main.zig

problem
The simulation runs up to 1 million users, so there are a lot of timelines. Those are inherently not upper bounded, so despite assuming a certain capacity, the timeline might need more memory, so the SMPAllocator is needed if the timeline needs to grow bigger thab the capacity.

The SMPAllocator uses a Slab of memory as a VMA for small (less than 64K) allocations, but when it needs more memory it mmaps a whole Slab for that specific object, and that adds a new VMA for that process.

Apparently, linux has a vm_max_alloc of arround 1 million VMA. If running the 100K topology with 12 workers, timelines keep betting bigger and bigger and the max VMA gets surpassed, and the process gets killed by plenty of RAM available on the machine.

I considered maybe that I was using the SMPAllocator wrong, but I generate one and all the threads share them, which is what the documentations says, as far as I understand.

solution
After tinkering to create an allocator to use a big mmap at the beginning of the simulation and share it by regions, I though that C might have an allocator that I can link, and I discovered the jemalloc allocator, which was exactly what I needed! By allocating different memory sized chunks it reduced the RAM ussge as well as avoided the max VMA problem.

question
Did I interpret wrongly the SMPAllocator/chosen a poorly design? Or its that my use case was not supported for this allocator? Should Zig consider new allocators in the std if yes?

Thank you for reading all of this :slight_smile:

1 Like

You could simply raise that limit.

Otherwise, I personally believe SmpAllocator to be flawed and it shouldn’t be used for critical applications that heavily rely on memory allocations. Maybe try linking in mimalloc?

I am currently working on a new version that will hopefully correct a bunch of issues (including this one!)

2 Likes

About raising the limit:

  1. I feel that linux defaults are normally reasonable. So changing them feels off to me but I don’t have much experience on wether if this is the right call. Regardless, this would just sweep the problem under the rug imo as will have to change it in every machine!
  2. Changing it requires sudo, which implies convincing my sysadmin that is a good idea, which requires time, resource I was short off xD

Never heard of it, but the nature of the solution is the same as using jemalloc: linking a c_allocator :))

Yayyy happy to hear that and looking forward to it! I thought of porting jemalloc to zig (just the allocation strategy part, basing it in page allocator as the DebugAllocators does) but I ended up just linking the c version.

Thank you for your answer :))

how are you allocating your timeline objects ? always one by one ? maybe consider batching the allocation ? if you store all the timeline into a “segmented list” does it help ?

I used to use a PaginatedMultiArrayList to store the Post of the simulation, so I know pagination (in fact, i implemented mine based in the same file you send ;)) but the problem with together allocation —as far as I understand it— is that every element in the page should be static if you want to allocate a buch together. Same for the page, the probles of the timelines is that they need to grow unboundedly as the simulation needs, so once the zone is bigger than 64KB, it will mmap them and bsck to the same problem.

Maybe we could make a big pool of memory that gets shared by some timelines, but that seems pretty complex to me. My first fix of this before dropping to use a c allocator was for the main thread to mmap a big ass chunck of memory and slipt it by workers, with as many VMA as workers. This worked, indeed, but used hella more memory that I was intending too :((