EDIT:
This post was flagged by community members as violating the AI policy of the Zig community forums. The BLUF is
Ziggit is by and for humans. We don’t want AI participation
I’m a human. I’m based in North London, UK, and trying to use the Zig compiler to build Grafana (written in Go) on one architecture for other architectures. Grafana currently doesn’t allow CGo dependencies, because we can’t afford to add the complexity of running multiple architectures to our build pipeline, but with Zig we could probably allow some C dependencies finally, such as Duck DB.
Continuing to review the AI policy at ziggit.dev/guidelines#AI then:
Do not use AI to generate code, and present it untested
I have used AI to generate some code here, but I have tested it, and found everything here to be correct.
Throwing in bad code just obfuscates the process even more
I’m not “throwing in bad code” here IMO - this is a solid reproduction case that I’ve built, to progress the issue at zig cc creating significantly slower binaries than GCC and clang in certain edge cases · Issue #16704 · ziglang/zig · GitHub - because I don’t have permissions to post to that issue directly myself.
Permitted uses of AI:
Translation: …
I’m new to Zig, and new to compiling C programs honestly. I did use AI to write the code that I wanted to run, but I ran it all myself, and indeed spotted the issue and pursued it myself with AI’s help. AI didn’t discover this issue, I did. But I don’t know all of the terminology that this community uses, so I wonder if you could afford me the option to “translate” my plain english into Zig community english, as I have done below. Please? Thanks in advance.
End of Edit
Start Here
Hi team
Nice to meet you all. Thanks for building/maintaining Zig - it seems a powerful too.
I wanted to post onto ziglang/zig/issues/16704 but I don’t have permissions, so posting here instead, to try to progress that issue.
I was cross-compiling DuckDB to aarch64-linux-musl with zig c++, building three times changing only -O1 / -O2 / -O3. All three produced byte-identical objects — and each was a genuine ~4.5-minute compile, not a cache hit. After digging (the trail leads to Zig’s ReleaseFast → clang -O2 mapping), here’s the minimal, reproducible picture.
Environment: Zig 0.16.0 (Homebrew build, bundled LLVM/clang 21.1.8), macOS 15.7.7, Apple M3 Pro, target aarch64-linux-musl.
Minimal repro:
printf 'long f(const int*a,unsigned long n){long s=0;for(unsigned long i=0;i<n;i++)s+=a[i];return s;}\n' > t.cpp
for O in O0 O1 O2 O3 Os Oz; do zig c++ -target aarch64-linux-musl -$O -c t.cpp -o z-$O.o; done
shasum -a 256 z-*.o
zig c++ collapses into three buckets:
z-O0.o distinct
z-O1.o ┐
z-O2.o ├─ byte-identical
z-O3.o ┘
z-Os.o ┐
z-Oz.o ┘─ byte-identical
So it isn’t only -O3→-O2: -O1 is folded in too, i.e. there’s no way to get a genuine -O1 (lighter optimization / faster compiles) either. The same source through host Apple clang 17 keeps -O1 distinct from -O2, so the source is optimization-sensitive and the collapse is the zig wrapper.
The mechanism, traced in the Zig source:
-
src/main.zig— the.optimizecase ziglang/zig/blob/738d2be9d6b6ef3ff3559130c05159ef53336224/src/main.zig#L2161-L2174 ORs-O1/-O2/-O3/-O4/-Ofastinto the sameoptimize_mode = .ReleaseFast(only-Os/-Oz→ReleaseSmall,-O0/-Og→Debug). -
src/Compilation.zig—.ReleaseFast =>ziglang/zig/blob/738d2be9d6b6ef3ff3559130c05159ef53336224/src/Compilation.zig#L7054-L7060 then appends clang-O2, with the comment: “we pass -O2 rather than -O3 because … -O3 is safer for Zig code than it is for C code … the -O3 path has been tested less.”
So -O1/-O2/-O3 → ReleaseFast → clang -O2 is by design. (Permalinks pinned to master @ 738d2be; 0.16.0 isn’t separately tagged on GitHub, but the code is identical — it’s at src/main.zig:2231 / src/Compilation.zig:6658 in the zig-0.16.0.tar.xz source.)
I tried both workarounds suggested in ziglang/zig#16704, for the compile-to-object (.o) case:
- The
-Xclang -Ofast/-Xclang -O3idea (from a commenter in that thread) is a no-op here.-###shows why — both zig’s injected-O2and the appended-O3land on the-cc1line, and the earlier-O2wins:
"-O2" # injected by zig (early)
"-O3" # from -Xclang -O3 (later, ignored)
- The maintainer’s recommended
-cflags … --path does get through for object output:zig build-obj -cflags -O3 -- t.cppproduces an object that differs from the-cflags -O2one. But it only takes effect in Zig’s default Debug mode, whereZIG_VERBOSE_CC=1shows clang also gets-O0(overridden by the-O3),-fsanitize=undefined, and frame-pointer flags. Adding-OReleaseFastto drop those re-collapses the level back to-O2.
So neither gives a clean -O3-release object. (This is the open issue tracking the perf angle: ziglang/zig#16704.)
Questions:
-
The
Compilation.zigcomment says the-O3path “has been tested less” for C (which is whyReleaseFastpasses clang-O2). Is there any plan to test/support a real-O3path for C in future — i.e. couldzig cc -O3eventually emit clang-O3? -
In the meantime, is there a supported way to opt into real
-O3, accepting the “tested less” caveat? For object builds I couldn’t find a clean one:-Xclang -O3onzig cc -cis a no-op;zig build-obj -cflags -O3reaches clang but only in Debug mode (drags in-fsanitize=undefinedetc.); and-OReleaseFastre-collapses to-O2. -
Motivation, in case it’s relevant: DuckDB’s own releases build with
-O3(its CMakeCMAKE_CXX_FLAGS_RELEASE), so I’m trying to match that when compiling its amalgamation withzig c++. In practice, how much does-O2-vs--O3matter for a large C++ codebase like this — is parity with upstream’s-O3worth chasing, or is-O2genuinely fine? (#16704 suggests it can matter for hot loops.)
Thanks! Happy to run experiments or test patches.
