I’ve been working on a uxn/varvara emulator, to get more familiar with zig and specifically to experiment with comptime code. The uxn design has 32 main instructions but each of those has 8 versions (eg. u8 or u16 values, the working stack or the return stack, etc.), and the instructions are encoded in a single byte where 3 bits represent the 8 flag combinations and the remaining 5 bits are the instruction. The C reference uxn code uses macros to abstract over the different versions and I wanted to see what that would look like using comptime. The result was not as succinct as the C code but I think it is easier to read.
This code switches on the 5 bit instruction code (I was happy to discover that zig knows when I’ve covered all the cases of a u5 in a switch). And I only have to list the 32 instructions rather than the full 256 instruction variants. It’s not shown here but op is a comptime parameter.
switch (op) {
// immediate op codes
0x20, 0x40, 0x60 => jxi(e, op),
0x80, 0xa0, 0xc0, 0xe0 => lit(e, flags),
// trucate to 5 bits == op & 0x1f
else => switch (@as(u5, @truncate(op))) {
0x00 => return false, // BRK
0x01 => inc(e, flags),
0x02 => pop(e, flags),
0x03 => nip(e, flags),
0x04 => swp(e, flags),
0x05 => rot(e, flags),
0x06 => dup(e, flags),
0x07 => ovr(e, flags),
0x08 => equ(e, flags),
0x09 => neq(e, flags),
0x0a => gth(e, flags),
0x0b => lth(e, flags),
0x0c => jmp(e, flags),
0x0d => jcn(e, flags),
0x0e => jsr(e, flags),
0x0f => sth(e, flags),
0x10 => ldz(e, flags),
0x11 => stz(e, flags),
0x12 => ldr(e, flags),
0x13 => str(e, flags),
0x14 => lda(e, flags),
0x15 => sta(e, flags),
0x16 => dei(e, flags),
0x17 => deo(e, flags),
0x18 => add(e, flags),
0x19 => sub(e, flags),
0x1a => mul(e, flags),
0x1b => div(e, flags),
0x1c => andOp(e, flags),
0x1d => ora(e, flags),
0x1e => eor(e, flags),
0x1f => sft(e, flags),
},
}
Then each instruction function takes the flags argument as a comptime argument, which means that we get separate compiled versions of each instruction depending on flags. I write one generic version and the compiler creates the 8 specific versions.
The pop and push functions are also comptime, so they push/pop u8 or u16 values depending on flags. This is generic code that I can understand.
fn ovr(e: *Evaluation, comptime flags: u8) void {
const b = e.pop(flags);
const a = e.pop(flags);
e.push(flags, a);
e.push(flags, b);
e.push(flags, a);
}
fn jmp(e: *Evaluation, comptime flags: u8) void {
if (comptime short(flags)) {
e.pc = e.pop(flags);
} else {
e.pc = e.pc +% relative(e.pop(flags));
}
}
Initially I executed a sequence of instructions with a while loop and inline else, which means that the switch expands to all 256 branches - that’s how I can treat the op as a comptime known value.
pub fn step(self: *Uxn) bool {
const op = self.readNextByte();
return switch (op) {
inline else => |o| self.execOp(o),
};
}
pub fn run(self: *Uxn) void {
while (self.step()) {}
}
Then later, I read about ‘labeled switch’ and was very happy to realise that I could get a tail-call/threaded-code version for almost no effort.
eval: switch (e.nextByte()) {
inline else => |op| {
if (execOp(&e, op)) continue :eval e.nextByte();
},
}
Looking at the assembly output, the evaluation is inlined completely, compiling to one massive jump table, with each of the 256 instructions tail-calling to the next instruction in the sequence - it’s a thing of beauty. Thanks to the neat design of uxn where stacks are 256 bytes and ram is 64k, I can index with u8 or u16 so the compiler removes all bounds checks.
I spent a long time implementing the graphics device using @floooh’s great sokol-zig - the time spent was entirely my fault, graphics APIs have moved on since I was last working with them.
Next should be the audio device but I’ve been putting that off, instead I watched the guide to data oriented design video and started experimenting with breaking up my structs. I found that keeping the machine’s scalar values, like the program counter and stack pointers, separate from the arrays used for stacks and ram made a big difference - almost 2x faster, but then a seemingly small, and I thought unrelated, change resulted in 3x slower. I should have waited until I had it feature complete, before profiling, because now every change I end up spending more time tweaking inline/noinline and looking at struct layouts to keep the performance.
So far, I’ve really enjoyed working with zig - I find it easier to read and write compared to C, and appreciate the friction and bumps that make me stop to consider the details of what I’m doing.