Issues with asm syntax

I have 2 problems:

1. How do tie output input registers?

I have a optimized dot function for sse4.1, but it manually allocates registers, thats less than ideal and will spawn a ton of vmovaps to copy from one register to another, LLVM has a syntax for it LLVM Language Reference Manual — LLVM 10 documentation that looks like =r,0 but zig can’t compile with that

pub inline fn dot_splat_sse41(v0: F32x4, v1: F32x4) F32x4 {
  return asm (
    \\dpps    $0xff, %[v1], %[v0]
    : [ret] "={xmm0}" (-> F32x4), // output
    : [v0] "{xmm0}" (v0), // inputs
      [v1] "x" (v1));
}

2. Multiple outputs and input constantness

The same link above states:

… It is not permitted for the asm to write to any input register or memory location (unless that input is tied to an output).

Now I implemented an optimized version of @reduce(.Min, v); for sse2, here I use the input v as a temp register, zig inline asm doesn’t allow multiple outputs, so what now?

Even if a choose to use a different register, I don’t know how do allocate a temp register.

pub inline fn f32x4_reduce_min(v: F32x4) f32 {
  if (sse2) {
    const out = asm (
      \\movaps   %[v],  %[r]
      \\unpckhpd %[r],  %[v]
      \\minps    %[v],  %[r]
      \\pshufd   $0x55, %[r], %[v]
      \\minps    %[v],  %[r]
      : [r] "=&x" (-> F32x4), // output
      : [v] "x" (v) // todo: LLVM prohibits from writing to %[v] because he isn't tied to an output
      : .{}
    );
return out[3];
  } else {
    @reduce(.Min, v); // todo:  NEON instruction for this
  }
}

Edit: fixed f32x4_reduce_min output type

I figureout 1, it turns out the link is talking about the LLMV IR syntax, gnu documentaion Simple Constraints (Using the GNU Compiler Collection (GCC)) doesn’t have a ,0 constraint only numbers that must be passed to the input like so:

pub inline fn dot_splat_sse41(v0: F32x4, v1: F32x4) F32x4 {
  return asm (
    \\dpps    $0xff, %[v1], %[v0]
    : [ret] "=x" (-> F32x4), // output
    : [v0] "0" (v0),
      [v1] "x" (v1));
}

Best i can suggest for 2 is to use an harcoded register for the second output. Explicitly clobber it in the main asm block, then use a second to copy out from the harcoded register.

In general forcing specific register may harm register allocation, but if you have a low density of such function it’s likely that LLVM can still generate optimal assembly