Runtime: Performance regression on Guid.Equals

Created on 11 Jul 2019  路  23Comments  路  Source: dotnet/runtime

When it's equal .NET Core is faster but when not, which happens probably more often, .NET Framework is faster.

BenchmarkDotNet=v0.11.5, OS=Windows 10.0.18362
Intel Core i7-4960X CPU 3.60GHz (Haswell), 1 CPU, 12 logical and 6 physical cores
.NET Core SDK=3.0.100-preview7-012593
  [Host] : .NET Core 3.0.0-preview7-27824-03 (CoreCLR 4.700.19.32302, CoreFX 4.700.19.32001), 64bit RyuJIT
  Clr    : .NET Framework 4.7.2 (CLR 4.0.30319.42000), 64bit RyuJIT-v4.8.3815.0
  Core   : .NET Core 3.0.0-preview7-27824-03 (CoreCLR 4.700.19.32302, CoreFX 4.700.19.32001), 64bit RyuJIT

| Method | Job | Runtime | Mean | Error | StdDev |
|------------ |----- |-------- |---------:|----------:|----------:|
| NoMatchGuid | Clr | Clr | 1.780 ns | 0.0016 ns | 0.0012 ns |
| NoMatchGuid | Core | Core | 2.279 ns | 0.0010 ns | 0.0008 ns |

[ClrJob]
[CoreJob]
[RPlotExporter]
public class GuidTest
{

    static byte[] FixedGUI = new byte[16] { 0x1, 0x1, 0x1, 0x2, 0x1, 0x1, 0x1, 0x2, 0x1, 0x1, 0x1, 0x2, 0x1, 0x1, 0x1, 0x2 };
    static byte[] FixedGUITwo = new byte[16] { 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5 };
    Guid OneGuid = new Guid(FixedGUI);
    Guid TwoGuid = new Guid(FixedGUITwo);

    [Benchmark]
    public bool NoMatchGuid() => OneGuid.Equals(TwoGuid);


}
area-System.Runtime tenet-performance

Most helpful comment

Codegen for Guid.Equals is the same in 2.2 and 3.0 (note this method is normally prejitted; in 3.0 the tiered version is more or less the same as the prejitted one). I don't see any regression locally.

BenchmarkDotNet=v0.11.5, OS=Windows 10.0.18917
Intel Core i7-4770HQ CPU 2.20GHz (Haswell), 1 CPU, 8 logical and 4 physical cores
.NET Core SDK=3.0.100-preview8-012981
  [Host]     : .NET Core 3.0.0-preview8-27910-02 (CoreCLR 4.700.19.35902, CoreFX 4.700.19.35911), 64bit RyuJIT
  Job-EAXLBN : .NET Framework 4.7.2 (CLR 4.0.30319.42000), 64bit RyuJIT-v4.8.3752.0
  Job-FXILZR : .NET Core 2.2.4 (CoreCLR 4.6.27521.02, CoreFX 4.6.27521.01), 64bit RyuJIT
  Job-TTFNJI : .NET Core 3.0.0-preview8-27910-02 (CoreCLR 4.700.19.35902, CoreFX 4.700.19.35911), 64bit RyuJIT


|                   Method | Runtime |     Toolchain |     Mean |     Error |    StdDev | Ratio | RatioSD |
|------------------------- |-------- |-------------- |---------:|----------:|----------:|------:|--------:|
|               EqualsSame |     Clr |        net472 | 7.135 ns | 0.0768 ns | 0.0718 ns |  1.00 |    0.00 |
|               EqualsSame |    Core | netcoreapp2.2 | 3.025 ns | 0.0383 ns | 0.0359 ns |  0.42 |    0.01 |
|               EqualsSame |    Core | netcoreapp3.0 | 2.348 ns | 0.0874 ns | 0.0897 ns |  0.33 |    0.01 |
|                          |         |               |          |           |           |       |         |
|  EqualsLastCharDifferent |     Clr |        net472 | 7.002 ns | 0.1750 ns | 0.2149 ns |  1.00 |    0.00 |
|  EqualsLastCharDifferent |    Core | netcoreapp2.2 | 2.717 ns | 0.0896 ns | 0.1165 ns |  0.39 |    0.01 |
|  EqualsLastCharDifferent |    Core | netcoreapp3.0 | 2.355 ns | 0.0461 ns | 0.0408 ns |  0.34 |    0.02 |
|                          |         |               |          |           |           |       |         |
| EqualsFirstCharDifferent |     Clr |        net472 | 6.956 ns | 0.0514 ns | 0.0456 ns |  1.00 |    0.00 |
| EqualsFirstCharDifferent |    Core | netcoreapp2.2 | 2.430 ns | 0.0886 ns | 0.1213 ns |  0.35 |    0.02 |
| EqualsFirstCharDifferent |    Core | netcoreapp3.0 | 2.764 ns | 0.0282 ns | 0.0264 ns |  0.40 |    0.00 |

So I'm still thinking the measured differences are some kind of artifact in the interaction of the benchmarking with HW.

Suggest we close this and keep an eye on the performance history of dotnet/performance#630, and maybe update that test also try first byte diff cases?

All 23 comments

@billwert

Also for this one @symbai would you mind comparing 2.1 or 2.2? If it's a regression its more interesting.

| Method | Job | Runtime | Toolchain | Mean | Error | StdDev |
|------------ |-------- |-------- |-------------- |---------:|----------:|----------:|
| NoMatchGuid | Default | Core | .NET Core 2.2 | 1.794 ns | 0.0702 ns | 0.0657 ns |
| NoMatchGuid | Clr | Clr | Default | 1.790 ns | 0.0107 ns | 0.0100 ns |
| NoMatchGuid | Core | Core | Default | 2.279 ns | 0.0008 ns | 0.0007 ns |

That seems odd, the code generated by .NET FX for Guid.Equals is worse than what .NET Core generates.

On my machine Core is faster than Clr:

BenchmarkDotNet=v0.11.5, OS=Windows 10.0.18932
Intel Core i5-4440 CPU 3.10GHz (Haswell), 1 CPU, 4 logical and 4 physical cores
.NET Core SDK=3.0.100-preview8-012936
  [Host] : .NET Core 3.0.0-preview8-27908-04 (CoreCLR 4.700.19.35801, CoreFX 4.700.19.35705), 64bit RyuJIT
  Clr    : .NET Framework 4.7.2 (CLR 4.0.30319.42000), 64bit RyuJIT-v4.8.3752.0
  Core   : .NET Core 3.0.0-preview8-27908-04 (CoreCLR 4.700.19.35801, CoreFX 4.700.19.35705), 64bit RyuJIT

| Method | Job | Runtime | Mean | Error | StdDev |
|------------ |----- |-------- |---------:|----------:|----------:|
| NoMatchGuid | Clr | Clr | 2.002 ns | 0.0206 ns | 0.0183 ns |
| NoMatchGuid | Core | Core | 1.906 ns | 0.0267 ns | 0.0249 ns |

I'm guessing this is another vectorization overhead vs first-byte different problem (similar to the String.Equals in dotnet/runtime#13057).

This is 6.5 vs 8 clock cycles. A difference in the GUID K byte would probably show 3.0 winning.

The perf test run seems to have HyperThreading on (more logical than physical cores), maybe this test just says "the hypercore was slower" (vs @mikedn's run which has it off (4 logical/4 physical)).

another vectorization overhead

Guid.Equals is not vectorized.

I am not able to reproduce the regression either. It is likely issue with code alignment or something similar.

I'm guessing this is another vectorization overhead vs first-byte different problem (similar to the String.Equals in dotnet/runtime#13057).

That's the problem, Guid.Equals is not even vectorized. vectorization could have been a source of surprises. But the change that was done to Equals is to compare the Guids as 4 integers:
```C#
return g._a == _a &&
Unsafe.Add(ref g._a, 1) == Unsafe.Add(ref _a, 1) &&
Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);

And there's nothing suspicious about the generated code:
```asm
G_M63787_IG01:
       0F1F440000           nop
G_M63787_IG02:
       8B02                 mov      eax, dword ptr [rdx]
       3B01                 cmp      eax, dword ptr [rcx]
       751D                 jne      SHORT G_M63787_IG04
       8B4204               mov      eax, dword ptr [rdx+4]
       3B4104               cmp      eax, dword ptr [rcx+4]
       7515                 jne      SHORT G_M63787_IG04
       8B4208               mov      eax, dword ptr [rdx+8]
       3B4108               cmp      eax, dword ptr [rcx+8]
       750D                 jne      SHORT G_M63787_IG04
       8B420C               mov      eax, dword ptr [rdx+12]
       3B410C               cmp      eax, dword ptr [rcx+12]
       0F94C0               sete     al
       0FB6C0               movzx    rax, al
G_M63787_IG03:
       C3                   ret
G_M63787_IG04:
       33C0                 xor      eax, eax
G_M63787_IG05:
       C3                   ret

Maybe some code alignment issue, changes in code size can have such effects. But that's basically luck.

Installed Preview 8, made the benchmark larger but still, CLR is faster. If you believe this is not an issue in .NET but has something to do with my machine etc feel free to close. I've got no idea just reporting what I've found.

BenchmarkDotNet=v0.11.5, OS=Windows 10.0.18362
Intel Core i7-4960X CPU 3.60GHz (Haswell), 1 CPU, 12 logical and 6 physical cores
.NET Core SDK=3.0.100-preview8-013015
  [Host] : .NET Core 3.0.0-preview8-27911-03 (CoreCLR 4.700.19.36002, CoreFX 4.700.19.36101), 64bit RyuJIT
  Clr    : .NET Framework 4.7.2 (CLR 4.0.30319.42000), 64bit RyuJIT-v4.8.3815.0
  Core   : .NET Core 3.0.0-preview8-27911-03 (CoreCLR 4.700.19.36002, CoreFX 4.700.19.36101), 64bit RyuJIT

| Method | Job | Runtime | Mean | Error | StdDev |
|-------- |----- |-------- |---------:|----------:|----------:|
| NoMatch | Clr | Clr | 13.76 ms | 0.1669 ms | 0.1393 ms |
| NoMatch | Core | Core | 16.05 ms | 0.1392 ms | 0.1234 ms |

[ClrJob]
[CoreJob]
[RPlotExporter]
public class ObjectTest
{
    const int Number = 5000000;
    Guid[] Objects = new Guid[Number];
    Guid[] ObjectsTwo = new Guid[Number];
    Random random = new Random();
    byte[] bytes = new byte[16];

    [GlobalSetup]
    public void Setup()
    {
        for (int i = 0; i < Number; i++)
        {
            random.NextBytes(bytes);
            Objects[i] = new Guid(bytes);
            random.NextBytes(bytes);
            ObjectsTwo[i] = new Guid(bytes);
            //Verify
            while (Objects[i].Equals(ObjectsTwo[i]))
            {
                random.NextBytes(bytes);
                ObjectsTwo[i] = new Guid(bytes);
            }
        }
    }

    [Benchmark]
    public bool NoMatch()
    {
        var result = false;
        for (int i = 0; i < Number; i++)
        {
            result = Objects[i].Equals(ObjectsTwo[i]);
        }
        return result;
    }

}

Results for the updated benchmark:

BenchmarkDotNet=v0.11.5, OS=Windows 10.0.18932
Intel Core i5-4440 CPU 3.10GHz (Haswell), 1 CPU, 4 logical and 4 physical cores
.NET Core SDK=3.0.100-preview8-012936
  [Host] : .NET Core 2.2.3 (CoreCLR 4.6.27414.05, CoreFX 4.6.27414.05), 64bit RyuJIT
  Clr    : .NET Framework 4.7.2 (CLR 4.0.30319.42000), 64bit RyuJIT-v4.8.3752.0
  Core   : .NET Core 2.2.3 (CoreCLR 4.6.27414.05, CoreFX 4.6.27414.05), 64bit RyuJIT

| Method | Job | Runtime | Mean | Error | StdDev |
|-------- |----- |-------- |---------:|----------:|----------:|
| NoMatch | Clr | Clr | 13.09 ms | 0.0379 ms | 0.0336 ms |
| NoMatch | Core | Core | 12.23 ms | 0.0649 ms | 0.0506 ms |

No idea what could be causing this difference, especially considering that we have processors from the same generation.

windows version is also different OS=Windows 10.0.18362 vs OS=Windows 10.0.18932. Just a guess.

Actually the above was mistakenly run with .NET Core 2.2. With .NET Core 3.0 there's basically no difference between Core and Clr:

BenchmarkDotNet=v0.11.5, OS=Windows 10.0.18932
Intel Core i5-4440 CPU 3.10GHz (Haswell), 1 CPU, 4 logical and 4 physical cores
.NET Core SDK=3.0.100-preview8-012936
  [Host] : .NET Core 3.0.0-preview8-27908-04 (CoreCLR 4.700.19.35801, CoreFX 4.700.19.35705), 64bit RyuJIT
  Clr    : .NET Framework 4.7.2 (CLR 4.0.30319.42000), 64bit RyuJIT-v4.8.3752.0
  Core   : .NET Core 3.0.0-preview8-27908-04 (CoreCLR 4.700.19.35801, CoreFX 4.700.19.35705), 64bit RyuJIT

| Method | Job | Runtime | Mean | Error | StdDev |
|-------- |----- |-------- |---------:|----------:|----------:|
| NoMatch | Clr | Clr | 13.11 ms | 0.0700 ms | 0.0655 ms |
| NoMatch | Core | Core | 13.17 ms | 0.0270 ms | 0.0225 ms |

OS, yeah, I suppose it could matter, though the versions seem to be rather close.
On the other hand the updated benchmark uses random values, that might make it more difficult to reproduce results.

Could be branch predictor aliasing-- the test is comparing random unequal GUIDs, so presumably the 4 branches in the GUID comparer are branching randomly 50/50. If any of those branches aliases vs some other live and usually predictable branch in the benchmark, it could cause perf to drop.

@adamsitnik do we have a way of getting CPI (cycles per instruction) data easily with BDN?

You might also want to control the random seed.

[edit: only the first 3 branches would be random 50/50... if they all compare equal the 4th branch must compare unequal...]

If the arrays are filled with the GUIDs from the initial version of the benchmark then Core does show a small slowdown:

BenchmarkDotNet=v0.11.5, OS=Windows 10.0.18932
Intel Core i5-4440 CPU 3.10GHz (Haswell), 1 CPU, 4 logical and 4 physical cores
.NET Core SDK=3.0.100-preview8-012936
  [Host] : .NET Core 3.0.0-preview8-27908-04 (CoreCLR 4.700.19.35801, CoreFX 4.700.19.35705), 64bit RyuJIT
  Clr    : .NET Framework 4.7.2 (CLR 4.0.30319.42000), 64bit RyuJIT-v4.8.3752.0
  Core   : .NET Core 3.0.0-preview8-27908-04 (CoreCLR 4.700.19.35801, CoreFX 4.700.19.35705), 64bit RyuJIT

| Method | Job | Runtime | Mean | Error | StdDev |
|-------- |----- |-------- |---------:|----------:|----------:|
| NoMatch | Clr | Clr | 13.00 ms | 0.0254 ms | 0.0237 ms |
| NoMatch | Core | Core | 13.21 ms | 0.0618 ms | 0.0578 ms |

The difference is very small but it seems to be consistent from run to run.

On the other hand, if all GUIDs are equal then Core is 2x faster:

| Method | Job | Runtime | Mean | Error | StdDev |
|-------- |----- |-------- |---------:|----------:|----------:|
| NoMatch | Clr | Clr | 32.83 ms | 0.1182 ms | 0.1106 ms |
| NoMatch | Core | Core | 16.29 ms | 0.0796 ms | 0.0744 ms |

For reference, the code generated by .NET FX is

00007FFCD7105430 0F 1F 44 00 00       nop         dword ptr [rax+rax]  
00007FFCD7105435 8B 02                mov         eax,dword ptr [rdx]  
00007FFCD7105437 3B 01                cmp         eax,dword ptr [rcx]  
00007FFCD7105439 75 64                jne         00007FFCD710549F  
00007FFCD710543B 48 0F BF 42 04       movsx       rax,word ptr [rdx+4]  
00007FFCD7105440 66 3B 41 04          cmp         ax,word ptr [rcx+4]  
00007FFCD7105444 75 59                jne         00007FFCD710549F  
00007FFCD7105446 48 0F BF 42 06       movsx       rax,word ptr [rdx+6]  
00007FFCD710544B 66 3B 41 06          cmp         ax,word ptr [rcx+6]  
00007FFCD710544F 75 4E                jne         00007FFCD710549F  
00007FFCD7105451 0F B6 42 08          movzx       eax,byte ptr [rdx+8]  
00007FFCD7105455 3A 41 08             cmp         al,byte ptr [rcx+8]  
00007FFCD7105458 75 45                jne         00007FFCD710549F  
00007FFCD710545A 0F B6 42 09          movzx       eax,byte ptr [rdx+9]  
00007FFCD710545E 3A 41 09             cmp         al,byte ptr [rcx+9]  
00007FFCD7105461 75 3C                jne         00007FFCD710549F  
00007FFCD7105463 0F B6 42 0A          movzx       eax,byte ptr [rdx+0Ah]  
00007FFCD7105467 3A 41 0A             cmp         al,byte ptr [rcx+0Ah]  
00007FFCD710546A 75 33                jne         00007FFCD710549F  
00007FFCD710546C 0F B6 42 0B          movzx       eax,byte ptr [rdx+0Bh]  
00007FFCD7105470 3A 41 0B             cmp         al,byte ptr [rcx+0Bh]  
00007FFCD7105473 75 2A                jne         00007FFCD710549F  
00007FFCD7105475 0F B6 42 0C          movzx       eax,byte ptr [rdx+0Ch]  
00007FFCD7105479 3A 41 0C             cmp         al,byte ptr [rcx+0Ch]  
00007FFCD710547C 75 21                jne         00007FFCD710549F  
00007FFCD710547E 0F B6 42 0D          movzx       eax,byte ptr [rdx+0Dh]  
00007FFCD7105482 3A 41 0D             cmp         al,byte ptr [rcx+0Dh]  
00007FFCD7105485 75 18                jne         00007FFCD710549F  
00007FFCD7105487 0F B6 42 0E          movzx       eax,byte ptr [rdx+0Eh]  
00007FFCD710548B 3A 41 0E             cmp         al,byte ptr [rcx+0Eh]  
00007FFCD710548E 75 0F                jne         00007FFCD710549F  
00007FFCD7105490 0F B6 42 0F          movzx       eax,byte ptr [rdx+0Fh]  
00007FFCD7105494 3A 41 0F             cmp         al,byte ptr [rcx+0Fh]  
00007FFCD7105497 75 06                jne         00007FFCD710549F  
00007FFCD7105499 B8 01 00 00 00       mov         eax,1  
00007FFCD710549E C3                   ret  
00007FFCD710549F 33 C0                xor         eax,eax  
00007FFCD71054A1 C3                   ret  

It's practically impossible for this code to be faster in any way for the 2 GUIDs in the initial example. In that case the very fist component of the GUID is different so the first branch will always be taken. In both code versions the target branch is more than 16 bytes away so another 16 byte block will have to be decoded anyway, this makes it improbable that there's a alignment induced issue.

.NET Core is faster when the guid is the same, a LOT faster. But there is a reproducable regression for .NET Core 2.2 vs .NET Core 3.0 for guids being different. No matter if the guid is hardcoded in the benchmark or random. I wonder where the difference come from when Mikedn says there is no room for improvements? (Not sure if it matters though)

| Method | Job | Runtime | Toolchain | Mean | Error | StdDev | Median |
|---------- |-------- |-------- |-------------- |---------:|---------:|---------:|---------:|
| Different | Default | Core | .NET Core 2.2 | 153.0 ms | 3.072 ms | 5.217 ms | 150.1 ms |
| Same | Default | Core | .NET Core 2.2 | 170.8 ms | 3.236 ms | 3.597 ms | 172.4 ms |
| Different | Clr | Clr | Default | 165.6 ms | 3.211 ms | 4.806 ms | 162.5 ms |
| Same | Clr | Clr | Default | 301.6 ms | 5.935 ms | 7.065 ms | 306.5 ms |
| Different | Core | Core | Default | 177.0 ms | 3.460 ms | 4.850 ms | 174.3 ms |
| Same | Core | Core | Default | 174.8 ms | 3.465 ms | 4.505 ms | 176.4 ms |

[ClrJob]
[CoreJob]
[RPlotExporter]
public class ObjectTest
{
    const int Number = 50000000;
    Guid[] Objects = new Guid[Number];
    Guid[] ObjectsTwo = new Guid[Number];
    Guid[] ObjectsSameAsOne = new Guid[Number];
    byte[] BytesOne = new byte[16] { 0x1, 0x1, 0x1, 0x2, 0x1, 0x1, 0x1, 0x2, 0x1, 0x1, 0x1, 0x2, 0x1, 0x1, 0x1, 0x2 };
    byte[] BytesTwo = new byte[16] { 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5, 0x5 };

    [GlobalSetup]
    public void Setup()
    {
        for (int i = 0; i < Number; i++)
        {
            Objects[i] = new Guid(BytesOne);
            ObjectsSameAsOne[i] = new Guid(BytesOne);
            ObjectsTwo[i] = new Guid(BytesTwo);
        }
    }

    [Benchmark]
    public bool Different()
    {
        var result = false;
        for (int i = 0; i < Number; i++)
        {
            result = Objects[i].Equals(ObjectsTwo[i]);
        }
        return result;
    }
    [Benchmark]
    public bool Same()
    {
        var result = false;
        for (int i = 0; i < Number; i++)
        {
            result = Objects[i].Equals(ObjectsSameAsOne[i]);
        }
        return result;
    }

}

While you're at it you may want to try this:
```C#
[Benchmark]
public bool NoMatchSSE()
{
var result = false;
for (int i = 0; i < Number; i++)
result = Eq(ref Objects[i], ObjectsTwo[i]);
return result;
}

static bool Eq(ref Guid x, Guid y)
{
return Sse2.MoveMask(Sse2.CompareEqual(Unsafe.As>(ref x), Unsafe.As>(ref y))) == 0xFFFF;
}
```
On my machine the SSE version is a bit faster than the current .NET Core 3.0 version:

| Method | Mean | Error | StdDev |
|----------- |---------:|----------:|----------:|
| NoMatch | 13.18 ms | 0.0370 ms | 0.0289 ms |
| NoMatchSSE | 11.34 ms | 0.0634 ms | 0.0593 ms |

You may also want to test using result |= instead of just result =.

I am unable to reproduce this regression on my box.

using BenchmarkDotNet.Attributes;
using BenchmarkDotNet.Running;
using System;

namespace GuidPerf
{
    class Program
    {
        static void Main(string[] args) => BenchmarkSwitcher.FromAssembly(typeof(Program).Assembly).Run(args);
    }

    public class Perf_Guid
    {
        const string guidStr = "a8a110d5-fc49-43c5-bf46-802db8f843ff";
        const string diffStr = "a8a110d5-fc49-43c5-bf46-802db8f843fe"; // last char is different

        private readonly Guid _guid = new Guid(guidStr);
        private readonly Guid _same = new Guid(guidStr);
        private readonly Guid _lastCharDifferent = new Guid(diffStr);

        [Benchmark]
        public bool EqualsSame() => _guid.Equals(_same);

        [Benchmark]
        public bool EqualsLastCharDifferent() => _guid.Equals(_lastCharDifferent);
    }
}
<Project Sdk="Microsoft.NET.Sdk">

  <PropertyGroup>
    <OutputType>Exe</OutputType>
    <TargetFrameworks>net472;netcoreapp2.2;netcoreapp3.0</TargetFrameworks>
  </PropertyGroup>

  <ItemGroup>
    <PackageReference Include="BenchmarkDotNet.Diagnostics.Windows" Version="0.11.5" />
  </ItemGroup>

</Project>
dotnet run -c Release -f netcoreapp2.2 --filter * --runtimes net472 netcoreapp2.2 netcoreapp3.0

```ini
BenchmarkDotNet=v0.11.5, OS=Windows 10.0.18362
Intel Xeon CPU E5-1650 v4 3.60GHz, 1 CPU, 12 logical and 6 physical cores
.NET Core SDK=3.0.100-preview7-012697
[Host] : .NET Core 2.2.0 (CoreCLR 4.6.27110.04, CoreFX 4.6.27110.04), 64bit RyuJIT
Job-GSIKNL : .NET Framework 4.7.2 (CLR 4.0.30319.42000), 64bit RyuJIT-v4.8.3801.0
Job-WIQUVU : .NET Core 2.2.0 (CoreCLR 4.6.27110.04, CoreFX 4.6.27110.04), 64bit RyuJIT
Job-LTHOEG : .NET Core 3.0.0-preview7-27826-20 (CoreCLR 4.700.19.32603, CoreFX 4.700.19.32613), 64bit RyuJIT
````

| Method | Runtime | Toolchain | Mean | Error | StdDev | Ratio |
|------------------------ |-------- |-------------- |---------:|----------:|----------:|------:|
| EqualsSame | Clr | net471 | 5.740 ns | 0.0125 ns | 0.0104 ns | 1.00 |
| EqualsSame | Core | netcoreapp2.2 | 2.109 ns | 0.0172 ns | 0.0152 ns | 0.37 |
| EqualsSame | Core | netcoreapp3.0 | 1.991 ns | 0.0062 ns | 0.0055 ns | 0.35 |
| | | | | | | |
| EqualsLastCharDifferent | Clr | net471 | 5.743 ns | 0.0193 ns | 0.0180 ns | 1.00 |
| EqualsLastCharDifferent | Core | netcoreapp2.2 | 2.165 ns | 0.0113 ns | 0.0094 ns | 0.38 |
| EqualsLastCharDifferent | Core | netcoreapp3.0 | 2.007 ns | 0.0132 ns | 0.0110 ns | 0.35 |

Funny, while we got a lot of different tests now the final meaning stays the same on my end.

My results of @mikedn suggestion to use Result |= using benchmark in https://github.com/dotnet/coreclr/issues/25644#issuecomment-510614537:

| Method | Job | Runtime | Toolchain | Mean | Error | StdDev |
|---------- |-------- |-------- |-------------- |---------:|----------:|----------:|
| Different | Default | Core | .NET Core 2.2 | 152.5 ms | 0.8862 ms | 0.7400 ms |
| Same | Default | Core | .NET Core 2.2 | 174.9 ms | 1.1041 ms | 0.9788 ms |
| Different | Clr | Clr | Default | 152.6 ms | 0.8500 ms | 0.7951 ms |
| Same | Clr | Clr | Default | 290.1 ms | 2.5175 ms | 2.2317 ms |
| Different | Core | Core | Default | 175.1 ms | 1.0461 ms | 0.9273 ms |
| Same | Core | Core | Default | 176.2 ms | 3.3806 ms | 2.9968 ms |

My result of @adamsitnik benchmark:

| Method | Job | Runtime | Toolchain | Mean | Error | StdDev | Median |
|------------------------ |-------- |-------- |-------------- |---------:|----------:|----------:|---------:|
| EqualsSame | Default | Core | .NET Core 2.2 | 2.328 ns | 0.0863 ns | 0.1292 ns | 2.257 ns |
| EqualsLastCharDifferent | Default | Core | .NET Core 2.2 | 2.293 ns | 0.0799 ns | 0.0784 ns | 2.256 ns |
| EqualsSame | Clr | Clr | Default | 5.561 ns | 0.1424 ns | 0.1332 ns | 5.487 ns |
| EqualsLastCharDifferent | Clr | Clr | Default | 5.731 ns | 0.1359 ns | 0.1271 ns | 5.685 ns |
| EqualsSame | Core | Core | Default | 2.601 ns | 0.0809 ns | 0.0757 ns | 2.561 ns |
| EqualsLastCharDifferent | Core | Core | Default | 2.577 ns | 0.0687 ns | 0.0643 ns | 2.557 ns |

My result of using @mikedn benchmark using SSE, which is faster but I cannot compare it against .NET Core 2.2 as APIs are missing:

| Method | Job | Runtime | Toolchain | Mean | Error | StdDev |
|---------- |-------- |-------- |-------------- |---------:|----------:|----------:|
| NoMatchSSE | Core | Core | Default | 128.9 ms | 1.0416 ms | 0.8698 ms |
| NoMatch | Core | Core | Default | 174.6 ms | 0.5818 ms | 0.4858 ms |
| MatchSSE | Core | Core | Default | 128.4 ms | 1.0798 ms | 0.9572 ms |
| Match | Core | Core | Default | 168.2 ms | 1.0851 ms | 0.8472 ms |

Looks like the SSE clearly positively outperforms everything. On .NET Core 3.0 only it gives a huge performance improvement in both cases.

I can't repro it with the bytes @adamsitnik uses, but I can repro it (using @adamsitnik's test) with the bytes @Symbai used.

From the analysis of the codegen it looks like things work as we expect and there's going to be cases where this is faster or slower. That's probably fine. I don't recall ever seeing Guid equality show up as a meaningful impact in a performance investigation. Unless the codegen folks (@AndyAyersMS) think there's a meaningful opportunity I don't think there's much to do.

Thanks for the report, @Symbai! It's an interesting case.

do we have a way of getting CPI (cycles per instruction) data easily with BDN?

@AndyAyersMS yes. You need to combine --profiler ETW with --counters counter1+counter2+counter3. To get the CPI I think that you should use TotalCycles+InstructionRetired (for some reason TotalCycles is always 0 on my PC)

To get the list of available counters you need to run following command from VS Command Prompt

tracelog.exe -profilesources Help

And run as admin on Windows :

dotnet run -c Release -f netcoreapp2.2 --filter * --counters InstructionRetired --profiler ETW --runtimes net472 netcoreapp2.2 netcoreapp3.0

Please keep in mind that the reported result contain the attached profiler overhead:

| Method | Runtime | Toolchain | Mean | Error | StdDev | Ratio | InstructionRetired/Op |
|------------------------ |-------- |-------------- |---------:|----------:|----------:|------:|----------------------:|
| EqualsSame | Clr | net472 | 6.130 ns | 0.0122 ns | 0.0102 ns | 1.00 | 52 |
| EqualsSame | Core | netcoreapp2.2 | 2.290 ns | 0.0454 ns | 0.0402 ns | 0.37 | 31 |
| EqualsSame | Core | netcoreapp3.0 | 2.219 ns | 0.0179 ns | 0.0168 ns | 0.36 | 31 |
| | | | | | | | |
| EqualsLastCharDifferent | Clr | net472 | 5.945 ns | 0.0156 ns | 0.0146 ns | 1.00 | 52 |
| EqualsLastCharDifferent | Core | netcoreapp2.2 | 2.290 ns | 0.0101 ns | 0.0095 ns | 0.39 | 31 |
| EqualsLastCharDifferent | Core | netcoreapp3.0 | 2.147 ns | 0.0633 ns | 0.0592 ns | 0.36 | 31 |

I'll take a look at the code paths involved just to be sure there's nothing unexpected here.

const string diffStr = "a8a110d5-fc49-43c5-bf46-802db8f843fe"; // last char is different

A difference in the last byte of the GUID won't ever show anything interesting because it results in all Equals code being run and there's very little chance that the huge .NET FX version will be anything other than slow. The interesting part is when the difference is in the first int of the GUID, that will result into an early out and then both .NET FX and .NET Core should have similar speed.

Codegen for Guid.Equals is the same in 2.2 and 3.0 (note this method is normally prejitted; in 3.0 the tiered version is more or less the same as the prejitted one). I don't see any regression locally.

BenchmarkDotNet=v0.11.5, OS=Windows 10.0.18917
Intel Core i7-4770HQ CPU 2.20GHz (Haswell), 1 CPU, 8 logical and 4 physical cores
.NET Core SDK=3.0.100-preview8-012981
  [Host]     : .NET Core 3.0.0-preview8-27910-02 (CoreCLR 4.700.19.35902, CoreFX 4.700.19.35911), 64bit RyuJIT
  Job-EAXLBN : .NET Framework 4.7.2 (CLR 4.0.30319.42000), 64bit RyuJIT-v4.8.3752.0
  Job-FXILZR : .NET Core 2.2.4 (CoreCLR 4.6.27521.02, CoreFX 4.6.27521.01), 64bit RyuJIT
  Job-TTFNJI : .NET Core 3.0.0-preview8-27910-02 (CoreCLR 4.700.19.35902, CoreFX 4.700.19.35911), 64bit RyuJIT


|                   Method | Runtime |     Toolchain |     Mean |     Error |    StdDev | Ratio | RatioSD |
|------------------------- |-------- |-------------- |---------:|----------:|----------:|------:|--------:|
|               EqualsSame |     Clr |        net472 | 7.135 ns | 0.0768 ns | 0.0718 ns |  1.00 |    0.00 |
|               EqualsSame |    Core | netcoreapp2.2 | 3.025 ns | 0.0383 ns | 0.0359 ns |  0.42 |    0.01 |
|               EqualsSame |    Core | netcoreapp3.0 | 2.348 ns | 0.0874 ns | 0.0897 ns |  0.33 |    0.01 |
|                          |         |               |          |           |           |       |         |
|  EqualsLastCharDifferent |     Clr |        net472 | 7.002 ns | 0.1750 ns | 0.2149 ns |  1.00 |    0.00 |
|  EqualsLastCharDifferent |    Core | netcoreapp2.2 | 2.717 ns | 0.0896 ns | 0.1165 ns |  0.39 |    0.01 |
|  EqualsLastCharDifferent |    Core | netcoreapp3.0 | 2.355 ns | 0.0461 ns | 0.0408 ns |  0.34 |    0.02 |
|                          |         |               |          |           |           |       |         |
| EqualsFirstCharDifferent |     Clr |        net472 | 6.956 ns | 0.0514 ns | 0.0456 ns |  1.00 |    0.00 |
| EqualsFirstCharDifferent |    Core | netcoreapp2.2 | 2.430 ns | 0.0886 ns | 0.1213 ns |  0.35 |    0.02 |
| EqualsFirstCharDifferent |    Core | netcoreapp3.0 | 2.764 ns | 0.0282 ns | 0.0264 ns |  0.40 |    0.00 |

So I'm still thinking the measured differences are some kind of artifact in the interaction of the benchmarking with HW.

Suggest we close this and keep an eye on the performance history of dotnet/performance#630, and maybe update that test also try first byte diff cases?

Was this page helpful?
0 / 5 - 0 ratings

Related issues

btecu picture btecu  路  3Comments

jamesqo picture jamesqo  路  3Comments

bencz picture bencz  路  3Comments

nalywa picture nalywa  路  3Comments

Timovzl picture Timovzl  路  3Comments