What Should We _Remove_ from Zig?

Personally I’d say anytype, I don’t like how opaque it is, at the very least the proposal with the |T| syntax would make it more readable because T could be named and therefore solve my biggest issue with anytype which is that on top of sometimes breaking zls, it also is very jarring and always requires to jump to definition, which breaks the flow. But i have to say now it’s much better since the WriterGate and the new Io interface, but with the previous writer/reader it was unreadable.

I’d like to remove the C type too, I don’t think they provide anything, especially [*c] I’m pretty use one of the many topics have made have been about this one and me not understanding it, because of how it coerce and what it express, and I don’t like it.

I also don’t like the implicit coercion of usize/isize I think it makes it exactly like C with the difference that at least the Zig compiler doesn’t silently coerce my size_t to whatever I’m assigning it to.

I’m not against the removal of ! at the same time I wonder if this could be solved better by basically a kind of fmt step, where zig can do something similar to the macro expansion and simply expand all your errors into custom error set for each function, but saying it out loud I’m not sure this is really useful.

Overall I’m pretty satisfied with current error handling, what would make it bearable is if at least the zig compiler was able to infer when an error has been handled and therefore doesn’t need to reapear in the function signature that handled it, at the same time I’d be worry because on one hand people being “lazy” with error handling means that you as the user often get the opportunity to choose directly how to handle the error from library code, which in my opinion is better than the library swallowing it for you.

I’d like to remove every single manage container, I think it encourage to think about things more.

I’d love zig to remove std.debug.assert, I think it would gain much more power as a builtin, and could tie very nicely into the build server protocol, because external tools could use those assertion builtin as a way to make static analysis better, (although this is probably possible by just seeing the undefined at the end but still).

Having said that I love the language, I think the team behind it is doing a rock solid job, and I’ll let them cook.

5 Likes

That’s a very convoluted way to express a simple requiremente: I want the bytes without an additional terminator.

Rust handles the distinction between those two types quite well one for ordinary string literals and another for C literals; I think Zig should benefit from a similar way and offer more control and explicitness in that regard, rather than storing an extra byte “just in case.”

const text = "test";

“test” has type *const [N:0]u8, which is a literal backed by four content bytes + a traling zero

I know we can quit that terminator using an array conversion or a comptime helper, etc.
But requesting a N bytes backing array should be similarly and so easy, concise and explicit.

The saving can be small for just one literal, but it can accumulate with many literals and matter in memory-constrained programs. Making this choice easier would fit Zig’s philosophy on explicitness and control over memory representation.

I can understand the convenience for C interoperability however i dont think that matter should dictate the default representation of Zig’s literals and in some aspects of the language. I would prefer ordinary literals without an implicit terminator with simple explicit way to construct a version with terminator when an interface requires one.

For literals knwon at compile time, this choice could be resolved during compilation.
Also interfaces requiring only a pointer and a length would not need additional terminator in the literal’s underlying storage.

This doesn’t make sense to me. The savings are O(N), where N is (more or less) the number of string literals present in your source code. Moreover, unless I’m wildly mistaken, the place where you are saving memory is in the binary of your program, where the combination “I have a lot of string literals” and “I am memory constrained” seems vanishingly unlikely to be correctly solved by “I will shave one byte off of my string literals” rather than “I will rewrite my program so as to put fewer string literals into my binary”.

3 Likes

Even more so with string literal deduplication. Now it’s “I will rewrite my program to put fewer unique string literals into my binary”

3 Likes

There’s sometimes when you cant do that. Programming covers many uses cases today.
Think about embedded systems or firmware, where the binary has to fit on a fixed flash budget.
In that context, saving space in the binary is useful, even if it doesnt reduce RAM usage.

“Use fewer literals” is not always equivalent. Strings can be required for command names, protocol identifiers, user facing messages, menu labels, diagnostic messages… removing those can remove functionality omiting unused terminators may preserves their contents.

For example if the firmware were 100 bytes over the budget maybe saving 200 bytes could matter, and the actual saving would need to be checked in the compiled output.

The idea is that even minimal savings can be significant when memory is limited.
Although Zig already allows for the creation of arrays without a terminating character, Im arguing that requesting that representation should be more straightforward.

I just think that i see this as a design trade-off worth improving: making unnecessary storage easier to avoid would align with Zig’s emphasis on explicitness and control over memory.

Girl, use less AI when you write your replies, please.

In such a situation, the best thing to do is clearly to create one (1) string literal in your binary by concatenating all the string literals in your source code and rewriting your source code to use slices of that literal. It’s technically suboptimal because you will still have one (1) null terminator.

3 Likes

I mostly write code for microcontrollers, so I can truly claim I write “memory constrained” programs. I have had to play the game of hunt-the-byte a few times to get things to run.

When I hear “memory constrained” I always want to ask “what type of memory is constrained?”.

Consider, for example, an ESP32-S3 chip, which is a fairly powerful IoT microcontroller: dual-core 240 MHZ CPU, 512KB of on-chip SRAM that serves as general-purpose memory. Sounds like a lot, until you find out WiFi / LWIP alone uses up half of it.

There is a problem, though: SRAM doesn’t retain it’s data when the chip is powered off, and having to re-program your IoT device every time it’s unplugged gets old after a while. To fix this, we wire up a NAND flash chip, and save the .text section (code and constant data, including string literals) onto it instead. These chips range from 4MB-16MB in size.

Then, when the chip powers on, the bootloader loads all the code from flash into SRAM before running it…

What if instead, we load only the most-commonly used code, then load the rest of the code and constant data through a small cache of SRAM as-and-when we need it? That leaves us with more SRAM to use on fun things like variables and DMA buffers.

As a result, when it’s time to go hunting for bytes, I don’t bother with static strings. It’s not because they don’t use a lot of memory (they do!) but because the type of memory they use - flash ROM - is not the one I care about. I’ve got enough flash memory to store my program three times over (and I usually do - it makes over-the-air updates simpler). The memory I run out of is SRAM, of which constant strings occupy precisely 0 bytes of.

P.S. this is one situation in which C strings have a slight advantage over slices - they can be stored in a single register or word, instead of two. They use slightly less stack memory as a result, which is well worth the extra null byte.

3 Likes

That’s actually a really good idea, but yeah im just going down the rabbit hole of the most optimal way. Seeing how far we can go reducing unnecesary storage and what the trade-offs would be. I dont speak native english sorry if its very google translator or whatever.

1 Like

Better yet, take it to its logical extreme and merge strings that have a common prefix & suffix.*

It sounds silly, but if (hypothetically) you’re desperate enough to scrape individual bytes off strings…

*I think it’s an NP-hard problem (to calculate the most efficient merge). But it wouldn’t be hard to brute-force, when the total size is so small.

1 Like

I mean, in a way, Zig already has a builtin for assertion, unreachable. Isn’t that what std.debug.assert already uses? I’ve always thought of it as a thin wrapper. It’s not much of a static analyzer if it can’t reason about a function call imho. You could even make assert an inlined function, so the IR would be inlined.

yeah there might be better solutions, I’d just think that it might be a good idea, plus I wonder if it would help with some hardware that have dedicated instructions for signing pointers, or variables, or to protect some parts of the code. Having said that assert is good enough, but I’m pretty sure it would be better as a builtin, for the convenience, to encourage good practice, and because it’s easy to see.

I think discussions about assert can be continued in one of those topics, I don’t want endless re-iteration of previous discussions:

4 Likes