Runtime: Performance issue with Vector<Int32> division operation

Created on 27 Jun 2018  路  11Comments  路  Source: dotnet/runtime

The division operation of the System.Numerics.Vector<T> type with System.Int32 values is slower than scalar division operation.

BenchmarkDotNet=v0.10.14, OS=Windows 10.0.17134
Intel Core i7-6800K CPU 3.40GHz (Skylake), 1 CPU, 12 logical and 6 physical cores
.NET Core SDK=2.1.301
  [Host]     : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT
  DefaultJob : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT

|             Method |       Mean |     Error |    StdDev |
|--------------------|------------|-----------|-----------|
| ScalarSingleDivide | 11.6010 ns | 0.0240 ns | 0.0213 ns |
| VectorSingleDivide |  0.9732 ns | 0.0083 ns | 0.0077 ns |
|  ScalarInt32Divide | 14.1547 ns | 0.0684 ns | 0.0606 ns |
|  VectorInt32Divide | 26.2876 ns | 0.1418 ns | 0.1327 ns |
00007ff8`f1631c00 ConsoleApp1.VectorBenchmarks.VectorSingleDivide()
            return _vector1Single / _vector2Single;
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`f1631c03 c4e17d104158    vmovupd ymm0,ymmword ptr [rcx+58h]
00007ff8`f1631c09 c4e17d104978    vmovupd ymm1,ymmword ptr [rcx+78h]
00007ff8`f1631c0f c4e17c5ec1      vdivps  ymm0,ymm0,ymm1
00007ff8`f1631c14 c4e17d1102      vmovupd ymmword ptr [rdx],ymm0
00007ff8`f1631c19 488bc2          mov     rax,rdx

00007ff8`f1631c00 ConsoleApp1.VectorBenchmarks.VectorInt32Divide()
            return _vector1Int32 / _vector2Int32;
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`f1631c08 488bd6          mov     rdx,rsi
00007ff8`f1631c0b c4e17d104118    vmovupd ymm0,ymmword ptr [rcx+18h]
00007ff8`f1631c11 c4e17d11442440  vmovupd ymmword ptr [rsp+40h],ymm0
00007ff8`f1631c18 c4e17d104138    vmovupd ymm0,ymmword ptr [rcx+38h]
00007ff8`f1631c1e c4e17d11442420  vmovupd ymmword ptr [rsp+20h],ymm0
00007ff8`f1631c25 488bca          mov     rcx,rdx
00007ff8`f1631c28 488d542440      lea     rdx,[rsp+40h]
00007ff8`f1631c2d 4c8d442420      lea     r8,[rsp+20h]
00007ff8`f1631c32 e8714afeff      call    00007ff8`f16166a8 (op_Division)
00007ff8`f1631c37 488bc6          mov     rax,rsi
  • Benchmarks project: ConsoleApp1.zip
  • System.Numerics.Vectors version: 4.5.0
area-System.Numerics

Most helpful comment

I think the documentation for the type Vector<T> can be significantly improved by providing information which operations on supported types are optimized by intrinsic functions. For example, something like this table can give a more clear picture at a glance:

| | + | - | * | / |
| --- | :---: | :---: | :---: | :---: |
sbyte | Yes | Yes | No | No
byte | Yes | Yes | No | No
short | Yes | Yes | Yes | No
ushort | Yes | Yes | No | No
int | Yes | Yes | Yes | No
uint | Yes | Yes | No | No
long | Yes | Yes | No | No
ulong | Yes | Yes | No | No
float | Yes | Yes | Yes | Yes
double | Yes | Yes | Yes | Yes

All 11 comments

@alexanderkozlenko, could you clarify why you closed the issue?

@tannergooding, fixed some issues in benchmarks project.

It may be another day or two before I can look at this more in depth. CC. @eerhardt in the meantime.

@alexanderkozlenko, it might help if you had a larger inner loop. Currently you are testing, essentially, a single instruction. This means that the function prologue/epilogue are actually having more impact on the measured time than the instruction you are trying to measure.

Generally speaking (for small tests like this), your [Benchmark] method should contain a for loop that executes the code you want to test several thousand times in a loop.

However, looking at the Vector<T> operator / implementation, it doesn't look like it is treated as an intrinsic.... @CarolEidt might have more context as to why.

100000 iterations per a benchmark method.

BenchmarkDotNet=v0.10.14, OS=Windows 10.0.17134
Intel Core i7-6800K CPU 3.40GHz (Skylake), 1 CPU, 12 logical and 6 physical cores
.NET Core SDK=2.1.301
  [Host]     : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT
  DefaultJob : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT

|             Method |      Mean |     Error |    StdDev |
|------------------- |-----------|-----------|-----------|
| ScalarSingleDivide |  70.44 us | 0.2372 us | 0.2218 us |
| VectorSingleDivide |  28.76 us | 0.0831 us | 0.0694 us |
|  ScalarInt32Divide |  95.22 us | 0.3153 us | 0.2949 us |
|  VectorInt32Divide | 271.72 us | 1.4552 us | 1.3612 us |

1 us: 1 Microsecond (0.000001 sec)
00007ff8`f3c01a50 ConsoleApp1.VectorBenchmarks.VectorSingleDivide()
                result = _vector1Single / _vector2Single;
                ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`f3c01a55 c4e17d104158    vmovupd ymm0,ymmword ptr [rcx+58h]
00007ff8`f3c01a5b c4e17d104978    vmovupd ymm1,ymmword ptr [rcx+78h]
00007ff8`f3c01a61 c4e17c5ec1      vdivps  ymm0,ymm0,ymm1

00007ff8`f3bf1b30 ConsoleApp1.VectorBenchmarks.ScalarInt32Divide()
                result = _vector1Int32 / _vector2Int32;
                ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`f3bf1a85 488d4c2460      lea     rcx,[rsp+60h]
00007ff8`f3bf1a8a c4e17d104618    vmovupd ymm0,ymmword ptr [rsi+18h]
00007ff8`f3bf1a90 c4e17d11442440  vmovupd ymmword ptr [rsp+40h],ymm0
00007ff8`f3bf1a97 c4e17d104638    vmovupd ymm0,ymmword ptr [rsi+38h]
00007ff8`f3bf1a9d c4e17d11442420  vmovupd ymmword ptr [rsp+20h],ymm0
00007ff8`f3bf1aa4 488d542440      lea     rdx,[rsp+40h]
00007ff8`f3bf1aa9 4c8d442420      lea     r8,[rsp+20h]
00007ff8`f3bf1aae e8f54bfeff      call    00007ff8`f3bd66a8 (op_Division)

The division operation of the System.Numerics.Vector type with System.Int32 values is slower than scalar division operation.

I'm not sure what the problem is. There is no integer vector division instruction and the benchmarks are anyway not comparable due to the scalar one dividing by a constant.

Here are the results with the same dividing approach for the scalar case.

BenchmarkDotNet=v0.10.14, OS=Windows 10.0.17134
Intel Core i7-6800K CPU 3.40GHz (Skylake), 1 CPU, 12 logical and 6 physical cores
.NET Core SDK=2.1.301
  [Host]     : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT
  DefaultJob : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT


|             Method |      Mean |     Error |    StdDev |
|--------------------|-----------|-----------|-----------|
| ScalarSingleDivide |  71.09 us | 0.0584 us | 0.0517 us |
| VectorSingleDivide |  27.89 us | 0.0376 us | 0.0333 us |
|  ScalarInt32Divide | 201.65 us | 0.2113 us | 0.1873 us |
|  VectorInt32Divide | 265.31 us | 0.5024 us | 0.4453 us |

1 us: 1 Microsecond (0.000001 sec)
00007ff8`5b0b1c20 ConsoleApp1.VectorBenchmarks.VectorSingleDivide()
                result = _vector1Single / _vector2Single;
                ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`5b0b1c25 c4e17d104160       vmovupd ymm0,ymmword ptr [rcx+60h]
00007ff8`5b0b1c2b c4e17d108980000000 vmovupd ymm1,ymmword ptr [rcx+80h]
00007ff8`5b0b1c34 c4e17c5ec1         vdivps  ymm0,ymm0,ymm1

00007ff8`5b0c1c20 ConsoleApp1.VectorBenchmarks.VectorInt32Divide()
                result = _vector1Int32 / _vector2Int32;
                ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`5b0c1c55 488d4c2460         lea     rcx,[rsp+60h]
00007ff8`5b0c1c5a c4e17d104620       vmovupd ymm0,ymmword ptr [rsi+20h]
00007ff8`5b0c1c60 c4e17d11442440     vmovupd ymmword ptr [rsp+40h],ymm0
00007ff8`5b0c1c67 c4e17d104640       vmovupd ymm0,ymmword ptr [rsi+40h]
00007ff8`5b0c1c6d c4e17d11442420     vmovupd ymmword ptr [rsp+20h],ymm0
00007ff8`5b0c1c74 488d542440         lea     rdx,[rsp+40h]
00007ff8`5b0c1c79 4c8d442420         lea     r8,[rsp+20h]
00007ff8`5b0c1c7e e8254afeff         call    00007ff8`5b0a66a8 (op_Division)

@mikedn, my expectation is that such abstraction like the Vector<T> type (which is used to achieve some performance benefits) must not silently provide with worse performance results. Even if there is no intrinsic for a particular operation on a particular type, the results must be almost the same (not 30% worse). As an option, it can throw an exception telling that the specified operation with the specified type is not eligible for optimization.

As an option, it can throw an exception telling that the specified operation with the specified type is not eligible for optimization.

Perhaps but for better or worse that ship has sailed, these operations can't suddenly start throwing such exceptions.

Even if there is no intrinsic for a particular operation on a particular type, the results must be almost the same (not 30% worse)

I suppose it may be possible to improve the current managed implementation, perhaps by special casing and unrolling the loops for certain types like Vector<int>. To what extent that would improve things I do not know but in the end it's probably always going to be slower due to the added cost of packing and unpacking vectors.

I think the documentation for the type Vector<T> can be significantly improved by providing information which operations on supported types are optimized by intrinsic functions. For example, something like this table can give a more clear picture at a glance:

| | + | - | * | / |
| --- | :---: | :---: | :---: | :---: |
sbyte | Yes | Yes | No | No
byte | Yes | Yes | No | No
short | Yes | Yes | Yes | No
ushort | Yes | Yes | No | No
int | Yes | Yes | Yes | No
uint | Yes | Yes | No | No
long | Yes | Yes | No | No
ulong | Yes | Yes | No | No
float | Yes | Yes | Yes | Yes
double | Yes | Yes | Yes | Yes

@alexanderkozlenko, if you open a new issue requesting the documentation update; I'll mark it as up-for-grabs

@tannergooding, created the issue dotnet/corefx#32184 for the documentation update.

Thanks.

Was this page helpful?
0 / 5 - 0 ratings

Related issues

Timovzl picture Timovzl  路  3Comments

nalywa picture nalywa  路  3Comments

btecu picture btecu  路  3Comments

EgorBo picture EgorBo  路  3Comments

aggieben picture aggieben  路  3Comments