The division operation of the System.Numerics.Vector<T> type with System.Int32 values is slower than scalar division operation.
BenchmarkDotNet=v0.10.14, OS=Windows 10.0.17134
Intel Core i7-6800K CPU 3.40GHz (Skylake), 1 CPU, 12 logical and 6 physical cores
.NET Core SDK=2.1.301
[Host] : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT
DefaultJob : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT
| Method | Mean | Error | StdDev |
|--------------------|------------|-----------|-----------|
| ScalarSingleDivide | 11.6010 ns | 0.0240 ns | 0.0213 ns |
| VectorSingleDivide | 0.9732 ns | 0.0083 ns | 0.0077 ns |
| ScalarInt32Divide | 14.1547 ns | 0.0684 ns | 0.0606 ns |
| VectorInt32Divide | 26.2876 ns | 0.1418 ns | 0.1327 ns |
00007ff8`f1631c00 ConsoleApp1.VectorBenchmarks.VectorSingleDivide()
return _vector1Single / _vector2Single;
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`f1631c03 c4e17d104158 vmovupd ymm0,ymmword ptr [rcx+58h]
00007ff8`f1631c09 c4e17d104978 vmovupd ymm1,ymmword ptr [rcx+78h]
00007ff8`f1631c0f c4e17c5ec1 vdivps ymm0,ymm0,ymm1
00007ff8`f1631c14 c4e17d1102 vmovupd ymmword ptr [rdx],ymm0
00007ff8`f1631c19 488bc2 mov rax,rdx
00007ff8`f1631c00 ConsoleApp1.VectorBenchmarks.VectorInt32Divide()
return _vector1Int32 / _vector2Int32;
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`f1631c08 488bd6 mov rdx,rsi
00007ff8`f1631c0b c4e17d104118 vmovupd ymm0,ymmword ptr [rcx+18h]
00007ff8`f1631c11 c4e17d11442440 vmovupd ymmword ptr [rsp+40h],ymm0
00007ff8`f1631c18 c4e17d104138 vmovupd ymm0,ymmword ptr [rcx+38h]
00007ff8`f1631c1e c4e17d11442420 vmovupd ymmword ptr [rsp+20h],ymm0
00007ff8`f1631c25 488bca mov rcx,rdx
00007ff8`f1631c28 488d542440 lea rdx,[rsp+40h]
00007ff8`f1631c2d 4c8d442420 lea r8,[rsp+20h]
00007ff8`f1631c32 e8714afeff call 00007ff8`f16166a8 (op_Division)
00007ff8`f1631c37 488bc6 mov rax,rsi
System.Numerics.Vectors version: 4.5.0@alexanderkozlenko, could you clarify why you closed the issue?
@tannergooding, fixed some issues in benchmarks project.
It may be another day or two before I can look at this more in depth. CC. @eerhardt in the meantime.
@alexanderkozlenko, it might help if you had a larger inner loop. Currently you are testing, essentially, a single instruction. This means that the function prologue/epilogue are actually having more impact on the measured time than the instruction you are trying to measure.
Generally speaking (for small tests like this), your [Benchmark] method should contain a for loop that executes the code you want to test several thousand times in a loop.
However, looking at the Vector<T> operator / implementation, it doesn't look like it is treated as an intrinsic.... @CarolEidt might have more context as to why.
100000 iterations per a benchmark method.
BenchmarkDotNet=v0.10.14, OS=Windows 10.0.17134
Intel Core i7-6800K CPU 3.40GHz (Skylake), 1 CPU, 12 logical and 6 physical cores
.NET Core SDK=2.1.301
[Host] : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT
DefaultJob : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT
| Method | Mean | Error | StdDev |
|------------------- |-----------|-----------|-----------|
| ScalarSingleDivide | 70.44 us | 0.2372 us | 0.2218 us |
| VectorSingleDivide | 28.76 us | 0.0831 us | 0.0694 us |
| ScalarInt32Divide | 95.22 us | 0.3153 us | 0.2949 us |
| VectorInt32Divide | 271.72 us | 1.4552 us | 1.3612 us |
1 us: 1 Microsecond (0.000001 sec)
00007ff8`f3c01a50 ConsoleApp1.VectorBenchmarks.VectorSingleDivide()
result = _vector1Single / _vector2Single;
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`f3c01a55 c4e17d104158 vmovupd ymm0,ymmword ptr [rcx+58h]
00007ff8`f3c01a5b c4e17d104978 vmovupd ymm1,ymmword ptr [rcx+78h]
00007ff8`f3c01a61 c4e17c5ec1 vdivps ymm0,ymm0,ymm1
00007ff8`f3bf1b30 ConsoleApp1.VectorBenchmarks.ScalarInt32Divide()
result = _vector1Int32 / _vector2Int32;
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`f3bf1a85 488d4c2460 lea rcx,[rsp+60h]
00007ff8`f3bf1a8a c4e17d104618 vmovupd ymm0,ymmword ptr [rsi+18h]
00007ff8`f3bf1a90 c4e17d11442440 vmovupd ymmword ptr [rsp+40h],ymm0
00007ff8`f3bf1a97 c4e17d104638 vmovupd ymm0,ymmword ptr [rsi+38h]
00007ff8`f3bf1a9d c4e17d11442420 vmovupd ymmword ptr [rsp+20h],ymm0
00007ff8`f3bf1aa4 488d542440 lea rdx,[rsp+40h]
00007ff8`f3bf1aa9 4c8d442420 lea r8,[rsp+20h]
00007ff8`f3bf1aae e8f54bfeff call 00007ff8`f3bd66a8 (op_Division)
The division operation of the System.Numerics.Vector
type with System.Int32 values is slower than scalar division operation.
I'm not sure what the problem is. There is no integer vector division instruction and the benchmarks are anyway not comparable due to the scalar one dividing by a constant.
Here are the results with the same dividing approach for the scalar case.
BenchmarkDotNet=v0.10.14, OS=Windows 10.0.17134
Intel Core i7-6800K CPU 3.40GHz (Skylake), 1 CPU, 12 logical and 6 physical cores
.NET Core SDK=2.1.301
[Host] : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT
DefaultJob : .NET Core 2.1.1 (CoreCLR 4.6.26606.02, CoreFX 4.6.26606.05), 64bit RyuJIT
| Method | Mean | Error | StdDev |
|--------------------|-----------|-----------|-----------|
| ScalarSingleDivide | 71.09 us | 0.0584 us | 0.0517 us |
| VectorSingleDivide | 27.89 us | 0.0376 us | 0.0333 us |
| ScalarInt32Divide | 201.65 us | 0.2113 us | 0.1873 us |
| VectorInt32Divide | 265.31 us | 0.5024 us | 0.4453 us |
1 us: 1 Microsecond (0.000001 sec)
00007ff8`5b0b1c20 ConsoleApp1.VectorBenchmarks.VectorSingleDivide()
result = _vector1Single / _vector2Single;
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`5b0b1c25 c4e17d104160 vmovupd ymm0,ymmword ptr [rcx+60h]
00007ff8`5b0b1c2b c4e17d108980000000 vmovupd ymm1,ymmword ptr [rcx+80h]
00007ff8`5b0b1c34 c4e17c5ec1 vdivps ymm0,ymm0,ymm1
00007ff8`5b0c1c20 ConsoleApp1.VectorBenchmarks.VectorInt32Divide()
result = _vector1Int32 / _vector2Int32;
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
00007ff8`5b0c1c55 488d4c2460 lea rcx,[rsp+60h]
00007ff8`5b0c1c5a c4e17d104620 vmovupd ymm0,ymmword ptr [rsi+20h]
00007ff8`5b0c1c60 c4e17d11442440 vmovupd ymmword ptr [rsp+40h],ymm0
00007ff8`5b0c1c67 c4e17d104640 vmovupd ymm0,ymmword ptr [rsi+40h]
00007ff8`5b0c1c6d c4e17d11442420 vmovupd ymmword ptr [rsp+20h],ymm0
00007ff8`5b0c1c74 488d542440 lea rdx,[rsp+40h]
00007ff8`5b0c1c79 4c8d442420 lea r8,[rsp+20h]
00007ff8`5b0c1c7e e8254afeff call 00007ff8`5b0a66a8 (op_Division)
@mikedn, my expectation is that such abstraction like the Vector<T> type (which is used to achieve some performance benefits) must not silently provide with worse performance results. Even if there is no intrinsic for a particular operation on a particular type, the results must be almost the same (not 30% worse). As an option, it can throw an exception telling that the specified operation with the specified type is not eligible for optimization.
As an option, it can throw an exception telling that the specified operation with the specified type is not eligible for optimization.
Perhaps but for better or worse that ship has sailed, these operations can't suddenly start throwing such exceptions.
Even if there is no intrinsic for a particular operation on a particular type, the results must be almost the same (not 30% worse)
I suppose it may be possible to improve the current managed implementation, perhaps by special casing and unrolling the loops for certain types like Vector<int>. To what extent that would improve things I do not know but in the end it's probably always going to be slower due to the added cost of packing and unpacking vectors.
I think the documentation for the type Vector<T> can be significantly improved by providing information which operations on supported types are optimized by intrinsic functions. For example, something like this table can give a more clear picture at a glance:
| | + | - | * | / |
| --- | :---: | :---: | :---: | :---: |
sbyte | Yes | Yes | No | No
byte | Yes | Yes | No | No
short | Yes | Yes | Yes | No
ushort | Yes | Yes | No | No
int | Yes | Yes | Yes | No
uint | Yes | Yes | No | No
long | Yes | Yes | No | No
ulong | Yes | Yes | No | No
float | Yes | Yes | Yes | Yes
double | Yes | Yes | Yes | Yes
@alexanderkozlenko, if you open a new issue requesting the documentation update; I'll mark it as up-for-grabs
@tannergooding, created the issue dotnet/corefx#32184 for the documentation update.
Thanks.
Most helpful comment
I think the documentation for the type
Vector<T>can be significantly improved by providing information which operations on supported types are optimized by intrinsic functions. For example, something like this table can give a more clear picture at a glance:| |
+|-|*|/|| --- | :---: | :---: | :---: | :---: |
sbyte| Yes | Yes | No | Nobyte| Yes | Yes | No | Noshort| Yes | Yes | Yes | Noushort| Yes | Yes | No | Noint| Yes | Yes | Yes | Nouint| Yes | Yes | No | Nolong| Yes | Yes | No | Noulong| Yes | Yes | No | Nofloat| Yes | Yes | Yes | Yesdouble| Yes | Yes | Yes | Yes