Runtime: Aligned load/store for Vector<T>

Created on 11 Apr 2019  路  5Comments  路  Source: dotnet/runtime

When using a collection such as Span<Vector<T>>, all loads/stores on the span are performed using unaligned instructions.

For sufficiently aware code, this might cause an unnecessary perf drawback. I think the "just works" current behavior is ideal, but we need an escape hatch for such code.

Proposed API

namespace System.Numerics
{
    public static class Vector
    {
        public static Vector<T> UnsafeLoadAligned<T>(in Vector<T> vector);
        public static void UnsafeStoreAligned<T>(out Vector<T> vector, Vector<T> value);
    }
}

This API should generate aligned instructions when available, e.g. movaps for SSE. The methods should not perform any sort of correctness checking outside of an Assert.

Related issue (perhaps dependency) dotnet/runtime#954

area-System.Numerics

Most helpful comment

@scalablecory. I don't think this is a good idea for System.Numerics.Vector<T>:

  • Unlike the HWIntrinsics, Vector<T> is considered to be a higher abstraction and purposefully doesn't expose the same level of granularity
  • On modern hardware (basically everything from the past 10 years); unaligned loads are as fast as aligned loads when the data is actually aligned.

    • The reason we expose both for the HWIntrinsic APIs is so that you can get the most efficient encoding on older hardware, while still having deterministic behavior on all hardware

    • When the data is unaligned, it should still be just as fast, except for reads crossing a cache-line or page boundary; where there will be a perf penalty

  • For non-stack allocated memory, the GC is free to move it as @GrabYourPitchforks indicated; so the LoadAligned instruction isn't usable outside of pinning (thus needing to take a T*)
  • For stack allocated memory, the JIT will already align the data in some cases and will emit aligned loads/stores

All 5 comments

If you have a GC-managed reference (_in_, _ref_, _out_), how do you know that the reference itself is properly aligned? The GC is free to move it around in a manner that doesn't guarantee 128-bit / 256-bit alignment unless you pin it, and if it's pinned then our APIs may as well accept pointers instead of GC-managed references.

@scalablecory. I don't think this is a good idea for System.Numerics.Vector<T>:

  • Unlike the HWIntrinsics, Vector<T> is considered to be a higher abstraction and purposefully doesn't expose the same level of granularity
  • On modern hardware (basically everything from the past 10 years); unaligned loads are as fast as aligned loads when the data is actually aligned.

    • The reason we expose both for the HWIntrinsic APIs is so that you can get the most efficient encoding on older hardware, while still having deterministic behavior on all hardware

    • When the data is unaligned, it should still be just as fast, except for reads crossing a cache-line or page boundary; where there will be a perf penalty

  • For non-stack allocated memory, the GC is free to move it as @GrabYourPitchforks indicated; so the LoadAligned instruction isn't usable outside of pinning (thus needing to take a T*)
  • For stack allocated memory, the JIT will already align the data in some cases and will emit aligned loads/stores

this might cause an unnecessary perf drawback

In addition to all of @tannergooding's and @GrabYourPitchforks's good comments, we also don't like to add public API based on speculation.

If you have a GC-managed reference (in, ref, out), how do you know that the reference itself is properly aligned? The GC is free to move it around in a manner that doesn't guarantee 128-bit / 256-bit alignment unless you pin it, and if it's pinned then our APIs may as well accept pointers instead of GC-managed references.

@GrabYourPitchforks indeed. My primary intent is to allow e.g:

Span<Vector<float>> span = GetSomePinnedAndAlignedMemory();
Vector<float> vec = Vector.UnsafeLoadAligned(span[0]);

Of course, you can't have a Vector<float>* -- but I agree using a float* parameter would be better to remove reduce subtly incorrect things one might do with the API.

Unlike the HWIntrinsics, Vector<T> is considered to be a higher abstraction and purposefully doesn't expose the same level of granularity ... On modern hardware (basically everything from the past 10 years); unaligned loads are as fast as aligned loads when the data is actually aligned.

@tannergooding I can appreciate not wanting to complicate an abstraction for hardware that is going out of style 馃憤.

we also don't like to add public API based on speculation.

@stephentoub that's a good stance. My speculation was about algorithms using the instructions being memory-bound or not, not about the benefits of aligned memory access.

Lets leave well-understood unaligned penalties for older hardware alone per @tannergooding's stance, and instead look at code generation:

Span<Vector<float>> dst, src;
dst[0] += src[0];

Lets say the user is not on a machine supporting AVX. You should get this code, which costs 12 bytes:

movups xmm0, [dst]
movups xmm1, [src]
addps xmm0, xmm1
movups [dst], xmm0

If instead the CLR had a way to know src is aligned, it could generate this, which costs 9 bytes:

movups xmm0, [dst]
addps xmm0, [src]
movups [dst], xmm0

It also removes any dependency effects on xmm1, which might save an additional dependency-breaking instruction down the line.

Of course, you can't have a Vector<float>*

Just as an FYI, this is now possible in C# 8 when the struct is statically verifiable to be unmanaged.

https://sharplab.io/#v2:EYLgtghgzgLgpgJwD4AEAMACFBGAdAOQFcxEBLAYygG4BYAKHpQGYNCA7KCAMziwCYMAYXoBvehglYWKACwYAsgAoAanHIwA9ggA8XADYaIMAHwAqDAAdV6rQEpxksXUkYAvg4kepWOfm0AVYxU1TR1A8ysQuwwAdwALRF5/DBBWNkg2CABzOAATLycXd2dJL2YfDAB5AKDrUJqIuuiRN3pXIA==

Was this page helpful?
0 / 5 - 0 ratings

Related issues

matty-hall picture matty-hall  路  3Comments

sahithreddyk picture sahithreddyk  路  3Comments

jamesqo picture jamesqo  路  3Comments

chunseoklee picture chunseoklee  路  3Comments

nalywa picture nalywa  路  3Comments