clang -march=native -ffp-contract=fast         -> 5.96046448e-08
clang -march=native -ffp-contract=off          -> 0
python, rounding a*b first as float32 does: 0 (matches the unfused build)
python, rounding once after the add, as a fused instruction does: 5.96046448e-08 (matches the fused build)
fused instruction in the assembly: ['vfmadd213ss\t%xmm2, %xmm1, %xmm0     # xmm0 = (xmm1 * xmm0) + xmm2']
