From https://github.com/dotnet/coreclr/pull/15069#discussion_r151746531 :
ReadOnlySpan.TryCopyTo is ~9.9%, but the Buffer.Memmove it's using for the actual copy is ~3.9%, so there's 6% in there that going to something other than the actual copy (the trace shows ~2.6% in the ReadOnlySpan.TryCopyTo body and ~3.4% in the Span.CopyTo body that TryCopyTo calls). If we're looking to optimize for such percentages, I'd prefer to see us start by looking at making TryCopyTo faster, as that'll accrue to many other scenarios.
cc @ahsonkhan @KrzysztofCwalina
My current thinking is to make an API Buffer.Memmove(ref byte dest, ref byte src, nuint len) that's a sister to Buffer.Memmove(byte* dest, byte* src, nuint len). This moves the pinning logic (and the associated stack setup overhead) from the caller to the p/invoke itself, so it effectively defers this logic until it's absolutely necessary.
Initial testing is promising. It appears to result in +65% throughput for small spans (< 64 bytes or so), with diminishing returns for longer spans. Even for 200-byte spans or so it's still around a +15% improvement. (The baseline is the existing Span.TryCopyTo code.)
If we could improve this, it would be a big win. CopyTo is used in lots of places.
Could you share the test that shows the improvements? I would like to see what exactly is being compared.
My current thinking is to make an API Buffer.Memmove(ref byte dest, ref byte src, nuint len)
We could do that with little risk for fast span just by having this be an internal API, right?
Yes; however, this optimization uses the same ByReference<T> trick under the covers described here. And as Jan points out, this isn't really the scenario ByReference<T> was intended for, so there's some nervousness as to whether it's reliable enough for this.
I expect to have perf numbers later today. Took a bad rebase and everything is terrible. -_-
Test application below. My build issues are now fixed and results are coming in.
using System;
using System.Diagnostics;
using System.Runtime.CompilerServices;
class Program
{
static void Main(string[] args)
{
const int ITER_COUNT = 250_000_000;
Span<byte> src = GetArray<byte>(64);
Span<byte> dest = new byte[32 * 1024];
Stopwatch sw = new Stopwatch();
while (true)
{
sw.Restart();
for (int i = 0; i < ITER_COUNT; i++)
{
// Codegen may inline the TryCopyTo method, but manual inspection
// shows it's not "cheating" e.g. by only comparing the source and
// destination lengths once then skipping the comparison on
// subsequent loop iterations.
if (!src.TryCopyTo(dest))
{
return;
}
}
// Console.WriteLine(sw.Elapsed);
// Console.WriteLine is currently misbehaving on my build, so this is a workaround.
foreach (var ch in sw.ElapsedMilliseconds.ToString())
{
Console.Write(ch);
}
Console.Write('\r');
Console.Write('\n');
}
}
[MethodImpl(MethodImplOptions.NoInlining)]
private static T[] GetArray<T>(int len)
{
return new T[len];
}
}
Perf measurements below. These measurements are for Span<T>.TryCopyTo, but similar measurements were found for ReadOnlySpan<T>.TryCopyTo.
Before and after times are in milliseconds and are for 250 million iterations of the inner test loop as described earlier. Test machine is Win10 x64 FCU. _Before_ is the average time taken for each test battery co complete before the code in the PR is applied; _After_ the code in the PR is applied. _Difference_ is the percentage difference in time taken (lower is better). _Throughput_ is the reciprocol of the difference and is the percentage improvement in calls per second (higher is better).
||Before (ms)|After (ms)|Difference|Throughput|
|---|---|---|---|---|
|4-byte span|1,823.94|1,138.45|-38%|+60%|
|64-byte span|1,789.90|1,038.81|-42%|+72%|
|256-byte span|4,406.38|3,722.88|-16%|+18%|
|64-object span|4,944.33|4,288.00|-13%|+15%|
Due to the way _Memmove_'s checks are ordered copying smaller-byte spans may be slightly slower than copying longer-byte spans, and this is reflected in the above tests. (I'm not considering changing this behavior as part of this investigation.)
As expected, the proposed changes show significant gains for smaller blittable spans with diminishing returns for longer blittable spans or spans of non-blittable types.
Most helpful comment
Perf measurements below. These measurements are for
Span<T>.TryCopyTo, but similar measurements were found forReadOnlySpan<T>.TryCopyTo.Before and after times are in milliseconds and are for 250 million iterations of the inner test loop as described earlier. Test machine is Win10 x64 FCU. _Before_ is the average time taken for each test battery co complete before the code in the PR is applied; _After_ the code in the PR is applied. _Difference_ is the percentage difference in time taken (lower is better). _Throughput_ is the reciprocol of the difference and is the percentage improvement in calls per second (higher is better).
||Before (ms)|After (ms)|Difference|Throughput|
|---|---|---|---|---|
|4-byte span|1,823.94|1,138.45|-38%|+60%|
|64-byte span|1,789.90|1,038.81|-42%|+72%|
|256-byte span|4,406.38|3,722.88|-16%|+18%|
|64-object span|4,944.33|4,288.00|-13%|+15%|
Due to the way _Memmove_'s checks are ordered copying smaller-byte spans may be slightly slower than copying longer-byte spans, and this is reflected in the above tests. (I'm not considering changing this behavior as part of this investigation.)
As expected, the proposed changes show significant gains for smaller blittable spans with diminishing returns for longer blittable spans or spans of non-blittable types.