Runtime: !Avx.Test{Z,C} suboptimal codegen

Created on 28 Nov 2018  路  5Comments  路  Source: dotnet/runtime

!Avx.TestZ and !Avx.TestC produce suboptimal code, can be shown with a simple repro:

```c#
using System.Runtime.CompilerServices;
using System.Runtime.Intrinsics;
using System.Runtime.Intrinsics.X86;

namespace ConsoleApp3
{
class Program
{
static int Main(string[] args)
{
Vector256 a = Avx.SetAllVector256(-42);
Vector256 b = Avx.SetAllVector256(42);

        bool res = Foo(a, b);

        return res ? 1 : 0;
    }

    [MethodImpl(MethodImplOptions.NoInlining)]
    private static bool Foo(Vector256<sbyte> a, Vector256<sbyte> b)
    {
        return !Avx.TestZ(a, b);
    }
}

}

results in (method `Foo`):
```asm
G_M7898_IG01:
       C5F877               vzeroupper 
       6690                 nop      

G_M7898_IG02:
       C4E17D10442408       vmovupd  ymm0, ymmword ptr[rsp+08H]
       C4E27D17442428       vptest   ymm0, ymmword ptr[rsp+28H]
       0F94C0               sete     al
       0FB6C0               movzx    rax, al
       85C0                 test     eax, eax
       0F94C0               sete     al
       0FB6C0               movzx    rax, al

G_M7898_IG03:
       C5F877               vzeroupper 
       C3                   ret      

Even withoud NoInlining this not ideal pattern can be seen.

Instead of the additional test-part, setne would be enough:

G_M7898_IG01:
       C5F877               vzeroupper
       6690                 nop

G_M7898_IG02:
       C4E17D10442408       vmovupd  ymm0, ymmword ptr[rsp+08H]
       C4E27D17442428       vptest   ymm0, ymmword ptr[rsp+28H]
-      0F94C0               sete     al
-      0FB6C0               movzx    rax, al
-      85C0                 test     eax, eax
-      0F94C0               sete     al
+      0F94C0               setne    al
       0FB6C0               movzx    rax, al

G_M7898_IG03:
       C5F877               vzeroupper
       C3                   ret

COMPlus_TieredCompilation=0 is set, so it is tier-1 code.

Codegen for the "positive" case (i.e. Avx.TestZ and Avx.TestC) is optimal.

category:cq
theme:intrinsics
skill-level:intermediate
cost:medium

area-CodeGen-coreclr optimization

Most helpful comment

I've verified that the change I have (https://github.com/mikedn/coreclr/commit/ffcd1a488e0a3ca05669ab8dc1deaa4e26ad2244) does generated the expected code:

G_M37424_IG01:
       C5F877               vzeroupper
       6690                 nop

G_M37424_IG02:
       C5FD10442408         vmovupd  ymm0, ymmword ptr[rsp+08H]
       C4E27D17442428       vptest   ymm0, ymmword ptr[rsp+28H]
       0F95C0               setne    al
       0FB6C0               movzx    rax, al

G_M37424_IG03:
       C5F877               vzeroupper
       C3                   ret

I should be able to finish this by the end of the week.

All 5 comments

More or less related to dotnet/runtime#9981. I could probably fix both at the same time.

@fiigii @CarolEidt @tannergooding

@mikedn May I ask the progress of https://github.com/dotnet/coreclr/issues/17073 fix?

@fiigii I have a fix somewhere in a branch but it depends on dotnet/coreclr#17733 getting merged.

I've verified that the change I have (https://github.com/mikedn/coreclr/commit/ffcd1a488e0a3ca05669ab8dc1deaa4e26ad2244) does generated the expected code:

G_M37424_IG01:
       C5F877               vzeroupper
       6690                 nop

G_M37424_IG02:
       C5FD10442408         vmovupd  ymm0, ymmword ptr[rsp+08H]
       C4E27D17442428       vptest   ymm0, ymmword ptr[rsp+28H]
       0F95C0               setne    al
       0FB6C0               movzx    rax, al

G_M37424_IG03:
       C5F877               vzeroupper
       C3                   ret

I should be able to finish this by the end of the week.

Was this page helpful?
0 / 5 - 0 ratings

Related issues

EgorBo picture EgorBo  路  3Comments

omajid picture omajid  路  3Comments

btecu picture btecu  路  3Comments

aggieben picture aggieben  路  3Comments

matty-hall picture matty-hall  路  3Comments