In the core of string.GetHashCode, there's some code that compiles conditionally for 32/64-bit, like so:
#if WIN32
// 32 bit machines.
int* pint = (int *)src;
int len = this.Length;
while (len > 2)
{
hash1 = ((hash1 << 5) + hash1 + (hash1 >> 27)) ^ pint[0];
hash2 = ((hash2 << 5) + hash2 + (hash2 >> 27)) ^ pint[1];
pint += 2;
len -= 4;
}
if (len > 0)
{
hash1 = ((hash1 << 5) + hash1 + (hash1 >> 27)) ^ pint[0];
}
#else
int c;
char *s = src;
while ((c = s[0]) != 0) {
hash1 = ((hash1 << 5) + hash1) ^ c;
c = s[1];
if (c == 0)
break;
hash2 = ((hash2 << 5) + hash2) ^ c;
s += 2;
}
#endif
Note that in the first implementation (for 32-bit) it uses the Length property to detect the end of the string, while in the second (for x64) it stops when it sees the first null character, which is rather strange since IIRC .NET strings allow you to have nulls in them, e.g.
Console.WriteLine(new string('\0', 1024).Length);
prints 1024. Is it intentional, then, for them to be ignored when hashing?
.NET Fiddle to demonstrate the bug
I ask this because I have a PR that could remove an unnecessary branch check from the method, but it might affect strings with embedded nulls.
See dotnet/coreclr#229
Summary of issue @mikedn links to Dictionary uses its own hashing rather than using string.GetHashCode
@mikedn Maybe we could fix this under an #if FEATURE_CORECLR? The new code doesn't look like it's being run (this value is always set to false), even for coreclr.
@benaadams Dictionary has this problem as well I think, since it uses this comparer which calls GetLegacyNonRandomizedHashCode which has the same problem.
Couldn't this null behavior make hash flooding attacks much easier? An attacker could send strings with str[0] == 0 causing all hash codes to be the same. That's easier than carefully crafting strings.
Most helpful comment
Couldn't this null behavior make hash flooding attacks much easier? An attacker could send strings with
str[0] == 0causing all hash codes to be the same. That's easier than carefully crafting strings.