Runtime: Is SpanHelpers simd-alignment safe?

Created on 22 Dec 2018  路  9Comments  路  Source: dotnet/runtime

https://github.com/dotnet/coreclr/blob/82e02f37564f83c3a7c90e5e28a652a39017701a/src/System.Private.CoreLib/shared/System/SpanHelpers.Byte.cs#L110-L114 (other places as well)
gets the count to register align the searchSpace, but it is not pinned so the GC can relocate it, thus the alignment is gone.

Is there any special handling in the VM / GC that I'm not aware of, or should this be rewritten to pin it?

The code still yields the correct results, but perf-wise it may be sub-optimal as potentially being not aligned correctly.
For tests there could be Debug.Assert added to check the alignment -- but this may result in flaky tests (depending on the non deterministic GC relocations).

area-System.Memory question

Most helpful comment

What is the benefit of the sequential alignment?

Alignment in the most common case

My concern is about cache line splits.

Impact of misalignment due to a thread suspending relocating GC is likely minor compared to the overhead of the threads being suspended?

All 9 comments

It's just a hint, which is why it also uses Unsafe.ReadUnaligned

What is the benefit of the sequential alignment then?

My concern is about cache line splits.

What is the benefit of the sequential alignment?

Alignment in the most common case

My concern is about cache line splits.

Impact of misalignment due to a thread suspending relocating GC is likely minor compared to the overhead of the threads being suspended?

Also pinning the memory to avoid the GC changing the alignment will likely cost more performance that misalignment would cause due to the extra bookkeeping the GC has to do and the fragementation it will cause to the GC heaps https://blogs.msdn.microsoft.com/maoni/2004/12/19/using-gc-efficiently-part-3/

As @benaadams said.

Thanks for clarifaction. So in the common case the data-access will be aligned. In the case the GC relocates the data, we don't care about the misalignment, as this is assumed to be rather rare.

A final question:

due to the extra bookkeeping the GC has to do and the fragementation it will cause to the GC heaps

I know the linked blog-post (anyway thx for sharing), but is this still true in contrast to https://github.com/dotnet/coreclr/pull/20275#discussion_r222882176?

Alignment in the most common case

@benaadams, I'm fairly certain this is not the case.

The GC does not currently support alignment above 8, so while you might get data that is 16/32-byte aligned, it will almost certainly not be the common case. I believe the stack (and some RVA statics) currently supports 16-byte alignment, but I'm not sure we currently have anything that guarantees that alignment, and that would still not guarantee alignment for Vector<T> when it is 32-bytes.

if (Vector.IsHardwareAccelerated && length >= Vector<byte>.Count * 2)

Given that the GC does not provide a way to guarantee alignment, that it is believed most machines will be newer than 6 years old (and will therefore have AVX/AVX2 support), and that the limit for using Vector<T> is at least 2 operations, this will almost certainly cause a cache-line split (unless we happen to get lucky and have 32-byte aligned data).

NOTE: I'm not saying we should pin, but there are potential perf optimizations to be had here still.... For example:

  • For sufficiently small data, it is likely more beneficial to use HWIntrinsics and only use the 128-bit instructions (mostly bits in Intel Optimization Manual, Chapter 12)
  • When using 256-bit instructions CPU downclocking can occur (Intel Optimization Manual 15.26) -- This section is specifically about Skylake, but it also discusses the higher impact on older micro-architectures
  • When using 256-bit instructions and the alignment is not known, using two 128-bit loads may prove more beneficial (also Chapter 12)
  • etc

Ah, I missed that the algorithm is computing the misalignment here..

It does alignment for the common case of no GC 馃榾

Was this page helpful?
0 / 5 - 0 ratings

Related issues

jkotas picture jkotas  路  3Comments

jzabroski picture jzabroski  路  3Comments

sahithreddyk picture sahithreddyk  路  3Comments

bencz picture bencz  路  3Comments

yahorsi picture yahorsi  路  3Comments