Std.mem.eql regression

Is this a valid regression? I just forwarded the post.

I don’t know if it’s a regression; but if it is, it should be explored why, instead of jumping to inline as a solution as that should be used sparingly due to its semantics.


I don’t think they would refuse a PR based on where the author works; it’s just not relevant.

I think any performance optimization also should come with a reproducible benchmark.
Also just inlining things has other costs, and may even have worse runtime performance in certain cases.
In this case I would want to see a simple example that clearly shows that it isn’t inlined in obvious cases. At which point it would probably be forwarded to llvm since it would likely affect more functions than just std.mem.eql.

…because it is pub fn without inline llvm doesn’t inline it as often as I want…

Correct me if I’m wrong, but whether a function is pub or not should have no impact on LLVM optimisations? Unlike C or C++, Zig doesn’t treat each file as a separate ā€˜compilation unit’, so there’s no need for pub functions to have global linkage. The effect he’s thinking of is export, I think.


People out there have some strange ideas about what Zig’s ā€œNo LLM policyā€ actually means, usually without having read it themselves.

I say that because Dmitriy himself appears to recognise the value in keeping your writing AI-free: GitHub - dmtrKovalenko/zlob: *Very* fast recursive file walking and globbing library for Zig, C, and Rust. With gitignore support. 100% POSIX compatible Ā· GitHub

P.S. No AI was used in the making of this README.md file thank you for reading it till the end.

A believe this is correct. There is actually a way to force llvm to inline function even if it’s not inline. @call has .always_inline option so he should just use that to test out if function been inlined really gives perf boost for his use case

IMO this shows serious limitation about how inlining works in LLVM.

let’s look at the code:

https://codeberg.org/ziglang/zig/src/commit/9276a56e3cee12a067f93d719a5e7f2762dd4603/lib/std/mem.zig#L799

it’s two simple guards and one for loops.
The guards should always be inlined but the for loop is probably better not inlined unless one of the two is small comptime know string.

There is a pass in LLVM that tries to break functions in smaller pieces and allow partial inlining but it’s not enabled by default, and Zig doesn’t enble it either.

When the inlining decision is taken, I don’t think LLVM looks at the arguments and leverage comptime info. It’s only once inlined that it can do so.

The first problem (splitting function) can be solved in usercode by manually splitting. Std already does this in some places.

I don’t know solutions for the second problem

just clarifying for others that zigs concept of comptime, and the optimisers’ (LLVM) are separate; ā€œconstant propagation/foldingā€ is the term usually used in optimiser land.