Runtime: Is it intentional that string.GetHashCode takes embedded nulls into account on 32-bit but not 64-bit platforms?

Created on 29 Apr 2016  路  4Comments  路  Source: dotnet/runtime

In the core of string.GetHashCode, there's some code that compiles conditionally for 32/64-bit, like so:

#if WIN32
                    // 32 bit machines.
                    int* pint = (int *)src;
                    int len = this.Length;
                    while (len > 2)
                    {
                        hash1 = ((hash1 << 5) + hash1 + (hash1 >> 27)) ^ pint[0];
                        hash2 = ((hash2 << 5) + hash2 + (hash2 >> 27)) ^ pint[1];
                        pint += 2;
                        len  -= 4;
                    }

                    if (len > 0)
                    {
                        hash1 = ((hash1 << 5) + hash1 + (hash1 >> 27)) ^ pint[0];
                    }
#else
                    int     c;
                    char *s = src;
                    while ((c = s[0]) != 0) {
                        hash1 = ((hash1 << 5) + hash1) ^ c;
                        c = s[1];
                        if (c == 0)
                            break;
                        hash2 = ((hash2 << 5) + hash2) ^ c;
                        s += 2;
                    }
#endif

Note that in the first implementation (for 32-bit) it uses the Length property to detect the end of the string, while in the second (for x64) it stops when it sees the first null character, which is rather strange since IIRC .NET strings allow you to have nulls in them, e.g.

Console.WriteLine(new string('\0', 1024).Length);

prints 1024. Is it intentional, then, for them to be ignored when hashing?

.NET Fiddle to demonstrate the bug

I ask this because I have a PR that could remove an unnecessary branch check from the method, but it might affect strings with embedded nulls.

area-System.Runtime question

Most helpful comment

Couldn't this null behavior make hash flooding attacks much easier? An attacker could send strings with str[0] == 0 causing all hash codes to be the same. That's easier than carefully crafting strings.

All 4 comments

See dotnet/coreclr#229

Summary of issue @mikedn links to Dictionary uses its own hashing rather than using string.GetHashCode

@mikedn Maybe we could fix this under an #if FEATURE_CORECLR? The new code doesn't look like it's being run (this value is always set to false), even for coreclr.

@benaadams Dictionary has this problem as well I think, since it uses this comparer which calls GetLegacyNonRandomizedHashCode which has the same problem.

Couldn't this null behavior make hash flooding attacks much easier? An attacker could send strings with str[0] == 0 causing all hash codes to be the same. That's easier than carefully crafting strings.

Was this page helpful?
0 / 5 - 0 ratings