Runtime: Consider updating 'BenchmarksGame' benchmarks to use HWIntrinsics

Created on 25 Oct 2019  路  22Comments  路  Source: dotnet/runtime

The "BenchmarkGames" for C# have been updated to use .NET Core 3.0: https://benchmarksgame-team.pages.debian.net/benchmarksgame/fastest/csharp.html

Given that .NET Core 3.0 now supports Hardware Intrinsics, it may be beneficial to updates these benchmarks to use them and see how much faster we can be.

We track our implementations of these toy benchmarks in the dotnet/performance repo: https://github.com/dotnet/performance/tree/master/src/benchmarks/micro/coreclr/BenchmarksGame and they would ideally also be contributed back to the Benchmark Games repo.

area-System.Runtime.Intrinsics up-for-grabs

All 22 comments

The contribution guide for the benchmark games repo is here: https://salsa.debian.org/benchmarksgame-team/benchmarksgame/blob/master/CONTRIBUTING.md

The information for how the programs are measured is here: https://benchmarksgame-team.pages.debian.net/benchmarksgame/how-programs-are-measured.html

CC. @danmosemsft, @adamsitnik

The benchmarks run on a Q6600 meaning that there's a few limitations as to what we can do:

We can only use up to SSE3, anything above it isn't supported on the CPU these benchmarks run on.
Apparently unaligned reads are very slow on a Q6600 so it's worth keeping it in mind.

There's definitely more things that are worth to keep in mind when creating a benchmark best suited for this.

Ideally we would write a benchmark that runs well on both modern machines with AVX support and older machines with only SSE3 support 馃槃

I agree, the benchmarks should support a variety of hardware, however this definitely needs to be taken into account when writing the benchmarks (as that's the CPU that's gonna showcase Core's speed).

I wonder whether the maintainer is looking for donations for a new machine 馃槃

cc @AnthonyLloyd who has done a lot of work on the C# and F# benchmarks in the past and may have suggestions for opportunities.

Current doing this for SpectralNorm, it looks good. Machine only has Ssse3 which is a little unfortunate for 3-vectors like n-body.

@AnthonyLloyd i assume this is your submission them? Nice result!
https://benchmarksgame-team.pages.debian.net/benchmarksgame/program/spectralnorm-csharpcore-5.html

@KoziLord any interest in taking a shot?

@danmosemsft I've tried a couple benchmarks in the past however I didn't get any improvements worth submitting. There's people more qualified to this than me, however if I do end up with something worthwhile I'll make sure to submit it.

Looks like there are HWIntrinsics implementations for all applicable benchmarks now. n-body was the last I remember being very slow, and it's now in line with other language/platform implementations.

Thanks @saucecontrol .

Looking at the list we are slower than Java only on reverse-complement. Could that use intrinsics?

https://benchmarksgame-team.pages.debian.net/benchmarksgame/program/revcomp-csharpcore-6.html

Oh yeah, looks like there's a relatively new C++ version that uses a clever SIMD LUT implementation. It's quite a bit faster than the others, despite being single-threaded.

Any interest in trying it in C# ? 馃槈

I can give it a look next time I have a free day, unless @john-h-k beats me to it ;)

Also of interest is that the 'official' hardware was fairly recently upgraded from the very old Intel Q6600 to a slightly newer Intel i5-3330, meaning SSE4.1 and AVX are now fair game (still no AVX2). Could be an opportunity to improve on all the scores again.

Oh - interesting, I did not know about the hardware upgrade.

Perhaps Spectral Norm could benefit from the upgrade to AVX?

"nbody" is already using AVX -- although there is a comment in there about random crashes. It was contributed by "Derek Ziemba" - not sure whether he has a Github ID.

Mandelbrot is using Vector.

I'll take another look at these. I have a new thing CsCheck that will be useful and it makes a good demo.

Would also be nice if we didn't have to delegate the heavy lifting to native libs (gmp, pcre) on some of those tests.

Spectral Norm doesn't seem to benefit from AVX at least on my laptop:

2.6%[-1..+0] slower, sigma=6.1 (19 vs 79)

Though in some good news (but not HWIntrinsics) I've just submitted and think I have something like this for reverse-complement:

25.1%[-5..+6] faster, sigma=6.0 (36 vs 0)

@AnthonyLloyd nice -- I"ll take it. Maybe we'll compare well with Java on all of them, after that goes in. Thank you!

Was this page helpful?
0 / 5 - 0 ratings