Runtime: Support for Intel SHA extensions

Created on 26 Nov 2019  路  10Comments  路  Source: dotnet/runtime

Intel SHA instructions assist with hardware acceleration of the SHA-1 and SHA-256 hash algorithms.

Current Ryzen processors that support these instructions can reach SHA-256 speeds of around 2 GB/s, roughly 4x faster than not using SHA instructions.

Support for these instructions via .NET Core's hardware intrinsics would be useful in SHA-heavy workloads.

Proposed API

namespace System.Runtime.Intrinsics.X86
{
    public class Sha1
    {
        /// <summary>
        /// __m128i _mm_sha1msg1_epu32 (__m128i a, __m128i b)
        /// SHA1MSG1 xmm, xmm/m128
        /// </summary>
        public static Vector128<byte> MessageSchedule1(Vector128<byte> a, Vector128<byte> b) => MessageSchedule1(a, b);

        /// <summary>
        /// __m128i _mm_sha1msg2_epu32 (__m128i a, __m128i b)
        /// SHA1MSG2 xmm, xmm/m128
        /// </summary>
        public static Vector128<byte> MessageSchedule2(Vector128<byte> a, Vector128<byte> b) => MessageSchedule2(a, b);

        /// <summary>
        /// __m128i _mm_sha1nexte_epu32 (__m128i a, __m128i b)
        /// SHA1NEXTE xmm, xmm/m128
        /// </summary>
        public static Vector128<byte> NextE(Vector128<byte> a, Vector128<byte> b) => NextE(a, b);

        /// <summary>
        /// __m128i _mm_sha1rnds4_epu32 (__m128i a, __m128i b, const int func)
        /// SHA1RNDS4 xmm, xmm/m128, imm8
        /// </summary>
        public static Vector128<byte> Rounds(Vector128<byte> state1, Vector128<byte> state2, byte func) => Rounds(state1, state2, func);
    }

    public class Sha256
    {
        /// <summary>
        /// __m128i _mm_sha256msg1_epu32 (__m128i a, __m128i b)
        /// SHA256MSG1 xmm, xmm/m128
        /// </summary>
        public static Vector128<byte> MessageSchedule1(Vector128<byte> a, Vector128<byte> b) => MessageSchedule1(a, b);

        /// <summary>
        /// __m128i _mm_sha256msg2_epu32 (__m128i a, __m128i b)
        /// SHA256MSG2 xmm, xmm/m128
        /// </summary>
        public static Vector128<byte> MessageSchedule2(Vector128<byte> a, Vector128<byte> b) => MessageSchedule2(a, b);

        /// <summary>
        /// __m128i _mm_sha256rnds2_epu32 (__m128i a, __m128i b, __m128i k)
        /// SHA256RNDS2 xmm, xmm/m128, &lt;XMM0&gt;
        /// </summary>
        public static Vector128<byte> Rounds(Vector128<byte> state1, Vector128<byte> state2, Vector128<byte> message) => Rounds(state1, state2, message);
    }
}

The naming of the functions and parameters still need work.

api-approved arch-x64 area-System.Runtime.Intrinsics

Most helpful comment

I agree that we should make these public for symmetry with AES, but I would strongly implore any developer __not__ to use these APIs unless you really, really, _really_ know what you're doing. Normally, writing secure assembly-level cryptography involves exercising tight control over register allocation, stack spilling, and knowledge about where your thread might be interrupted. The .NET Framework doesn't give developers such low-level control over program compilation and execution, and it's enormously easy to shoot yourself in the foot by making an algorithm which _appears_ to be faster but which in actuality inadvertently leaks secrets.

All 10 comments

@Thealexbarney thanks for the proposal. Our API review process is detailed here: https://github.com/dotnet/runtime/blob/master/docs/project/api-review-process.md

Could you please try to rewrite the original post to explicitly call out the APIs you would like to be added. We have a "good example" of a complete proposal here: https://github.com/dotnet/corefx/issues/271

Namely, we are looking for something like the Proposed API section at a minimum, but adding additional sections can help with the overall review process.

In this particular case it looks like there are 7 APIs that would be split between System.Runtime.Intrinsics.X86.Sha1 and System.Runtime.Intrinsics.X86.Sha256

Support for these instructions via .NET Core's hardware intrinsics would be useful in SHA-heavy workloads.

I'm curious what the use case is for these intrinisics? The SHA implementations provided by OpenSSL, CNG, etc should already be using these intrinisics. For example, https://github.com/openssl/openssl/blob/3c957bcd54d097167e53660fa100aa1ba85d63b5/crypto/sha/asm/sha256-mb-x86_64.pl#L539

Realistically no one should be using these intrinsics and they should only be using the known good implementations from locations like System.Security.

However, we've also said that we don't see a reason to block developers from having access to them as other languages (C, C++, Rust, etc) expose the corresponding intrinsics and there may be valid use cases for non-cryptographic scenarios (for example, a fast checksum of a downloaded web file without taking a dependency on a third party, not installed by default, library).

I'm curious what the use case is for these intrinisics? The SHA implementations provided by OpenSSL, CNG, etc should already be using these intrinisics.

The use case would be similar to the use case of AES-NI intrinsics.
Calling out to an outside implementation often has a performance impact, especially when working with small message sizes. Support for intrinsics would avoid that overhead by keeping that computation in-process. I suppose it also gives developers more freedom in creating their own implementations, although I don't know what benefit that would provide.

I agree that we should make these public for symmetry with AES, but I would strongly implore any developer __not__ to use these APIs unless you really, really, _really_ know what you're doing. Normally, writing secure assembly-level cryptography involves exercising tight control over register allocation, stack spilling, and knowledge about where your thread might be interrupted. The .NET Framework doesn't give developers such low-level control over program compilation and execution, and it's enormously easy to shoot yourself in the foot by making an algorithm which _appears_ to be faster but which in actuality inadvertently leaks secrets.

I've updated the top comment with what the API would look like. The naming of the functions and parameters still need work.

And as a side-note, my current use cases for both AES and SHA intrinsics don't need to be completely secure. Access to intrinsics has greatly sped up AES modes that aren't supported by .NET, close to the performance hand-written assembly would get. XTS mode is ~3-4x faster and doesn't require any tricks to reduce overhead from the crypto API calls.

@tannergooding is there anything missing from the proposal? If not I suppose we can mark it ready for review.

@tannergooding do you expect this to be reviewed? should it be marked milestone 5.0?

I've marked it future. We don't have any critical asks for this and the general recommendation is to use the official crypto APIs, so this will get reviewed as part of the general backlog.

Video

  • Looks good as proposed:

    • Add IsSupported

    • Change the Round methods to indicate number of rounds

    • Change parameters to match C++ naming

  • We should have an analyzer that flags usages, just like for the AES types. @tannergooding, please file the request and label it with code-analyzer.

```C#
namespace System.Runtime.Intrinsics.X86
{
public class Sha1
{
public static bool IsSupported { get; }
public static Vector128 MessageSchedule1(Vector128 a, Vector128 b);
public static Vector128 MessageSchedule2(Vector128 a, Vector128 b);
public static Vector128 NextE(Vector128 a, Vector128 b);
public static Vector128 FourRounds(Vector128 a, Vector128 b, byte func);
}

public class Sha256
{
    public static bool IsSupported { get; }
    public static Vector128<byte> MessageSchedule1(Vector128<byte> a, Vector128<byte> b);
    public static Vector128<byte> MessageSchedule2(Vector128<byte> a, Vector128<byte> b);
    public static Vector128<byte> TwoRounds(Vector128<byte> a, Vector128<byte> b, Vector128<byte> k);
}

}
```

Was this page helpful?
0 / 5 - 0 ratings

Related issues

chunseoklee picture chunseoklee  路  3Comments

jkotas picture jkotas  路  3Comments

bencz picture bencz  路  3Comments

jchannon picture jchannon  路  3Comments

EgorBo picture EgorBo  路  3Comments