Runtime: Is PInvoke or marshalling a large byte[] slower on MSVC?

Created on 5 Jun 2020  路  7Comments  路  Source: dotnet/runtime

General Summary

I'm using PInvoke in .NET Core 3.1 to pass a large byte[] into native code. The code is relatively simple and the DllImported method takes only blittable arguments. When I check the data transfer rate across various platforms, Ubuntu Linux x86_64 and macOS 10.15 XCode 11.5 both achieve on the order of 3.8 GB/sec. However Windows 10 Visual Studio 2019 16.6 x64 manages only about 50% of that. I know that this benchmark is comparing apples and oranges as the hardware of the desktop computers is not completely identical (though pretty similar) and the OS is different.

However, the difference in performance is not easily explained by the subtle differences in hardware, nor would I believe that the OS accounts for such a drastic difference. Furthermore, the fact that Windows manages relatively close to 50% of the rate makes me wonder if there is anything specific happening on Windows compared to the other platforms? Does Windows require more marshalling for byte[] -> unsigned char[] mapping? Are there other factors that come into play?

Environment

I've observed this behavior in .NET Core 3.1 on Ubuntu Linux x86_64, macOS 10.15 XCode 11.5, and Windows 10 Visual Studio 2019 16.6 x64.

On Linux, I use:

#> dotnet --info
.NET Core SDK (reflecting any global.json):
 Version:   3.1.201
 Commit:    b1768b4ae7

Runtime Environment:
 OS Name:     ubuntu
 OS Version:  18.04
 OS Platform: Linux
 RID:         ubuntu.18.04-x64
 Base Path:   /usr/share/dotnet/sdk/3.1.201/

Host (useful for support):
  Version: 3.1.3
  Commit:  4a9f85e9f8

.NET Core SDKs installed:
  3.1.201 [/usr/share/dotnet/sdk]

.NET Core runtimes installed:
  Microsoft.AspNetCore.App 3.1.3 [/usr/share/dotnet/shared/Microsoft.AspNetCore.App]
  Microsoft.NETCore.App 3.1.3 [/usr/share/dotnet/shared/Microsoft.NETCore.App]

To install additional .NET Core runtimes or SDKs:
  https://aka.ms/dotnet-download
area-Interop-coreclr question

Most helpful comment

Can you provide the benchmark you're using? P/Invoke of byte[] shouldn't be slower on Windows, all things being equal.

All 7 comments

Can you provide the benchmark you're using? P/Invoke of byte[] shouldn't be slower on Windows, all things being equal.

Here is the code: https://github.com/BioDataAnalysis/dotnet-runtime-37552

If you want to build and run it, please use the cmake instructions in InteropTestNative/CMakeLists.txt first. It should build without any dependencies, just like:

mkdir -p InteropTestNative/obj && \
cd InteropTestNative/obj/ && \
cmake .. -DCMAKE_INSTALL_PREFIX=../bin && \
make && \
make install

Afterwards you should be able to run dotnet test in NativeBinding.Tests. If there are problems resolving InteropTestNative you may need to tweak the path in NativeWrapper.cs.

@emmenlau There is a lot in that example. Can you point to the specific test or P/Invoke that is interesting/slow?

byte[] -> unsigned char[] mapping?

That is an unusual mapping if that is managed code. The char in managed code is always 2-bytes regardless of platform. If that is a native mapping the marshalling should be the same regardless of platform. Knowing the specific P/Invoke that is causing the confusion or represents the bottle neck would be helpful.

@AaronRobinsonMSFT sorry, I tried to keep the full benchmark usable, but it would have been good to highlight the interesting parts. I'll try to outline here.

This is the relevant Dllimport:
https://github.com/BioDataAnalysis/dotnet-runtime-37552/blob/70eca84c98c7d4a452d1c188796253820b78242b/NativeBinding.Tests/src/NativeWrapper.cs#L124

The managed code has signature:
```C#
[DllImport(cNativeImportName, EntryPoint = "test_bool_bytearray", CharSet = CharSet.Ansi, CallingConvention = CallingConvention.Cdecl, ExactSpelling = true)]
public static extern bool test_bool_bytearray(uint aWidth, uint aHeight, [In] byte[] aData);


Here is the corresponding native code:
https://github.com/BioDataAnalysis/dotnet-runtime-37552/blob/70eca84c98c7d4a452d1c188796253820b78242b/InteropTestNative/src/InteropTestNativeExport.hh#L40

The native code has signature:
```C++
INTEROPTESTNATIVE_EXPORT bool test_bool_bytearray(const unsigned int aWidth,
    const unsigned int aHeight,
    const unsigned char* aData);

I understand that byte[] should be an array unsigned 8 bit integer, and the corresponding native type is unsigned char*, is that correct?

I have significantly reduced the complexity of the sample code, in case that makes it more useful.

Also, I've seen that we get the same drop in performance when running a debug-version of the code. I do not know why Windows would run a debug version, so I do not believe this is the cause, but I will try to follow this up.

I understand that byte[] should be an array unsigned 8 bit integer, and the corresponding native type is unsigned char*, is that correct?

That is correct.

Building anything as Debug disables most if not all JIT optimizations so running performance related tests should always be done in Release. Measuring these kind of scenarios in .NET is hard given the nature of the system and complexity in how each interacts with the other. It is recommended to collect data under a profiler (e.g. PerfView or BPF) or to use BenchmarkDotNet which does a good job of avoiding many of the worse mistakes in these kind of measurements.

@AaronRobinsonMSFT Thanks for the quick reply. I'm not intentionally performing any benchmarks as debug, and I went to some effort to ensure that debug and release builds will go into dedicated output directories. I was just bringing this up because it was an interesting find, that debug performance on Linux/macOS matches release performance on Windows. I will ensure that there is no way that dotnet would mix up debug/release paths on Windows.

Was this page helpful?
0 / 5 - 0 ratings