Type compatibility of number operations make no sense

Why are signed and unsigned integers inherently incompatible with each other? I find myself adding say a usize and an i8 together often. Why are these two incompatible?
Like sure, you could argue “oh, its cuz if the usize value is less than 255 it’ll underflow”, and the compiler wants to protect you from that happening. But the compiler is not doing anything but adding more work to do.

I need to add a whole @as(T, @intCast()). The compiler is not protecting against anything here. I’d argue this also makes it more unreadable. It would be better if it was easier to type, but no, I have to surround the expression in two parenthesis, type in @ and some intrinsic names, it’s not easy to type.

And I have to do this INCREDIBLY often. Maybe my choice of types is poor, and I don’t know how to do it in idiomatic Zig, but the this weird behavior just makes me put signed types everywhere for convenience, even if I know they will never be negative, just to dodge having to put more @as and @intCast.

You use unsigned values when you know they will never be negative, obviously, and when something will be negative, obviously you use signed. Like, say I have a car with a position, which will always be positive, and maybe some wheel of the car has a negative position relative to the car. If I have to get the world position of the wheel, I need to add signed and unsigned positions to each other. I don’t care if maybe it underflows, I can handle that when it happens. But @intCast will not help in handling that.

And also, why are result types not propagated to the operands of an operation? This is also a source of annoyance for me. If

const a: u16 = 9;
const b: u8 = @intCast(a);

works, why shouldn’t

const a: u16 = 9;
const b: u8 = 1 + @intCast(a);

work?

I have to do this instead, which just adds to the visual clutter and it’s hard to read anything.

const b: u8 = 1 + @as(u8, @intCast(a));

How do I solve these problems? I genuinely can’t think of a way to structure the types so that they lessen the need of these intrinsices. The fact that loops always have a usize capture does not help either.

1 Like

I think every developer would agree that software is supposed to do stuff.
It is therefore beneficial for that stuff to be defined, both in concept, and implementation. That way it is clear how to implement the software, and it is clear what the implementation can and can’t do.
Undefined behaviour is problematic because it means some part, or all, of the software cannot be guaranteed to always do what it is intended to do.

Where developers will disagree is the degree in which they care, which also depends on the specific project in question.

Zig is catered for cases where you do care about that a decent to a lot.
Zig doesn’t prevent you writing undefined (or what Zig calls illegal) behaviour, because doing so may incorrectly prevent some software/implementations.
Rather Zig guides away from it with friction, and makes it explicitly visible when needed.

Sure, this does mean more work on the developers part, that is the cost. Hopefully, ranged integers will be implemented which will greatly improve ergonomics around integers (this is planned, but has not been attempted yet).

A cast isn’t protection, it is opting out of protection, it is saying you know better. Zig will, in safe modes, insert checks to verify what you are telling it; but not all things are, or can be checked.

Being mindful of integer types does help a lot, another strategy is to expand the type before a section that would otherwise require many casts, and cast back afterwards.

Another thing to keep in mind, Zig will coerce (implicitly cast) when the conversion has no chance of being unsafe.
For example: a smaller integer to a larger one, or the easily missed, unsigned → signed (where the signed integer is 1 bit larger).

then don’t use @intCast, there are the @add/subWithOverflow built-in’s, and std.math has some nice wrappers

it can change the result due to overflows, there are plenty of real cases where you want one or the other behaviour, so Zig makes you explicitly choose.

5 Likes

You must be new to Zig (or Rust for that matter which is much stricter) :wink:

Many of those “problems” exist because Zig doesn’t have C-style integer promotion (which is its own can of worms though - GCC and Clang have the -Wsign-conversion warning to get an idea how “sign-mismatch-polluted” a C or C++ code base is). There’s also a pretty long list of Zig proposals to ease those issues, so the problems are definitely known, but the question is where the sweet spot between robustness and convenience should be.

My general advice would be though to not mix signed and unsigned types in the first place, no matter what language (e.g. my rule of thumb both in C and Zig is to use “almost always signed”, unsigned integers should only be used for bit twiddling and modulo-math).

I fully agree about the “usize problem” though. The usize type shouldn’t be baked as much into the language as it currently is, or at the very least array sizes, indices and loop counters should be changed to isize.

PS:

just makes me put signed types everywhere for convenience, even if I know they will never be negative

This is exactly the right thing to do. AFAIK that whole problematic idea of sizes being unsigned goes back to C++'s size_t, and Stroustrup himself later argued that making size_t unsigned was a “bad mistake” (see “Subscripts and sizes should be signed”: https://open-std.org/jtc1/sc22/wg21/docs/papers/2019/p1428r0.pdf)

9 Likes

First why it happens, then how I personally deal with this problem:

I have had that thought many times myself, but have yet to be proven right.

In this case, it makes the code both easier to read and easier to compile. Consider the following C code. Does it have a bug?

int8_t aa = 200;
uint8_t bb = 1;

cc = aa + bb;

Trick question, you can’t know for sure without looking up the declaration of cc:

  • If it’s unsigned, the answer is 201 as expected.
  • If it’s signed, you get integer overflow and anything could happen.

With zig, those two cases look different:

const aa: i8 = 200;
const bb: u8 = 1;

cc = @as(u8, @intCast(aa)) + bb; // if cc: u8 (or larger)
cc = aa + @as(i8, @intCast(bb)); // if cc: i8 (or larger)

Use the wrong one and you get a compiler error.

This makes it easier for people reading your code (e.g. yourself in 2 weeks) because you can see what steps the computation is taking without having to look-up the type of cc.

For the same reason, it’s easier on the compiler - it can process the RHS without knowing what cc is, which makes compilation simpler, and allows more work to be done in parallel.


How to deal with it

When I find myself doing lots of a particular little conversion:

var len: u32 = 100;
const dd: i8 = 10;

len += itou32(dd);

...
/// Signed -> Unsigned
inline fn itou32(a: i32) u32 {
    return @intCast(a);
}
/// Unsigned -> Signed
inline fn utoi32(a: u32) i32 {
    return @intCast(a);
}

Names are just an example; call it whatever makes sense to you. If someone else reads the code and doesn’t know what it is, they can just search the file and find it’s definition.

3 Likes

Not really, I’ve had this problem for quite some time, I just haven’t had the energy to create an account and a post about it. I started using Zig like near end of last year I think, so I guess not very new?

I actually haven’t had as much of a problem with Rust, of course its strict but many of the errors I could solve without too much struggle.

Yeah, I have read about that, but I’m still not sure why its a “bad mistake”.

1 Like

I guess there are many different opinions, but C’s mixed-sign artithmetic only feels “natural” because it first converts everything to a signed integer (the only related historical mistake is that int is stuck at 32-bits even on 64-bit CPUs).

Tbh, I sometimes wonder whether carrying ‘signedness’ with the type was a mistake to begin with in programming language design.

It might have made a lot of sense before 2’s complement encoding was common. With 2’s complement, most integer operations are sign-agnostic anyway, and the problem is mostly reduced to ‘signed vs unsigned sign-extension’ (e.g. whether widening an integer requires the topmost bit to be replicated or zeroed). E.g. languages could have a single sign-agnostic “Schroedinger’s integer-type” that only becomes signed or unsigned “when looking into the box” (same discussion could be had for the width, e.g. the number of bits in an integer only matters for loading from and storing to memory, but all operations on the integer can happen on the ‘natural’ CPU register width).

7 Likes

In some cases - where I need a u32, but need a lot of calculations or initializations with i32 I invented this thing which can bypass @as. Just access i or u, whatever we need.

pub const UInt = packed union {
    u: u32,
    i: i32,

    pub fn init(i: i32) UInt {
        return .{ .u = @intCast(i) };
    }
};

To bypass @intFromEnum and @enumFromInt for (simple) enums I invented this:

pub const MyEnum = packed union {
    pub const count: usize = 4;

    pub const E = enum(u8) {
        zero = 0,
        one = 1,
        two = 2,
        three = 3,
    };

    e: E,
    u: u8,
};

I also wrote int and float conversion functions to bypass the terrible @as(i8, @intCast(bb)).

I agree with Floooh that the heavily baked in usize is not ideal. We need typed loops, for example.

2 Likes

It’s worth noting that is an area of ongoing ergonomic improvement, 0.16 made a number of changes in regards to type coercion largely around float/int interactions. Look for This is part of a larger effort to improve ergonomics for making video games in Zig..

As demonstrated above this can’t be extended to every single int conversion, but there are absolutely cases where Zig can improve and it’s heading in that direction already.

On a personal note I find having to cast ints a lot to be a clear demonstration I’m writing brittle code reliant on a high number of value range assumptions.

2 Likes

I agree with you that it’s annoying but it’s also how you can specify clearly your software, I’ve had many issues with C implicit promotion rules, and it’s honestly really annoying to have to be the compiler when reading code, and figure out, which operation gets promoted to what, and if it will overflow, or underflow.

In zig it’s explicit, which is nice, it’s verbose sure, but explicit. All i can say is to be patient, there are a lot of upcoming improvements in those areas, ranged integers, better result type propagation/inference, safer/more narrowing/widening when possible etc. Even for float I think it would be nice if they could select an implicit conversion of all floats to integers. There’s also upcoming support for different float, new types like Matrix, vec2/3/4/ etc. It’s probably going to be better in a few versions. In the meantime use comptime, or make small helpers.

1 Like

Many people encounter the same situation as you. I’m also looking forward to relevant improvement.

1 Like