Runtime: MS Office Addin using netcoreapp3.1 can crash PowerPoint

Created on 8 Jul 2020  路  14Comments  路  Source: dotnet/runtime

Description

  1. Create simple MS Office Addin using COM+ and netcoreapp3.1 for PowerPoint 32-bit.
  2. Register the comhost.dll generated for the assembly and register addin.
  3. Run PowerPoint 365 (Version 2005)
  4. Open COM Add-ins dialog and observe the custom addin is loaded
  5. Click the checkbox to disable the addin and click OK
  6. Open COM Add-ins dialog, enable the custom addin and click OK
  7. PowerPoint will crash.
Exception thrown at 0x00660066 in POWERPNT.EXE: 0xC0000005: Access violation executing location 0x00660066.

Configuration

  • Which version of .NET is the code running on? .NET Core 3.1.5
  • What OS and version, and what distro if applicable? Windows 10 2004 Enterprise x64, MS Office 365
  • What is the architecture (x64, x86, ARM, ARM64)? x86
  • Do you know whether it is specific to that configuration? I don't know

Other information

Visual Studio Output:

Exception thrown at 0x76159862 in POWERPNT.EXE: Microsoft C++ exception: Roaming::RoamingCacheException at memory location 0x243EEFCC.
Exception thrown at 0x76159862 in POWERPNT.EXE: Microsoft C++ exception: [rethrow] at memory location 0x00000000.
Exception thrown at 0x76159862 in POWERPNT.EXE: Microsoft C++ exception: Roaming::RoamingCacheException at memory location 0x243EEFCC.
Exception thrown at 0x76159862 in POWERPNT.EXE: Microsoft C++ exception: [rethrow] at memory location 0x00000000.
Exception thrown at 0x76159862 (KernelBase.dll) in POWERPNT.EXE: WinRT originate error - 0x83750007 : ''input': Value must not be empty.'.
Exception thrown at 0x76159862 (KernelBase.dll) in POWERPNT.EXE: WinRT originate error - 0x83750007 : ''input': Value must not be empty.'.
Exception thrown at 0x76159862 (KernelBase.dll) in POWERPNT.EXE: WinRT originate error - 0x83750007 : ''input': Value must not be empty.'.
Exception thrown at 0x76159862 (KernelBase.dll) in POWERPNT.EXE: WinRT originate error - 0x83750007 : ''input': Value must not be empty.'.
onecore\com\combase\catalog\catalog.cxx(2392)\combase.dll!7740B0F9: (caller: 7740B2EE) ReturnHr(15) tid(1dd0) 800401F3 Invalid class string
onecore\com\combase\catalog\catalog.cxx(2392)\combase.dll!7740B0F9: (caller: 7740B2EE) ReturnHr(16) tid(1dd0) 800401F3 Invalid class string
onecore\com\combase\catalog\catalog.cxx(2392)\combase.dll!7740B0F9: (caller: 7740B2EE) ReturnHr(17) tid(1dd0) 800401F3 Invalid class string
onecore\com\combase\catalog\catalog.cxx(2392)\combase.dll!7740B0F9: (caller: 7740B2EE) ReturnHr(18) tid(1dd0) 800401F3 Invalid class string
Exception thrown at 0x00660066 in POWERPNT.EXE: 0xC0000005: Access violation executing location 0x00660066.

Callstack:

>   MSO.DLL!05d0adab()  Unknown
    MSO.DLL![Frames below may be incorrect and/or missing, no symbols loaded for MSO.DLL]   Unknown
    MSO.DLL!05d0a391()  Unknown
    MSO.DLL!05d13948()  Unknown
    MSO.DLL!05d13901()  Unknown
    PPCORE.DLL!039e8c43()   Unknown
    MSO.DLL!05cf2e66()  Unknown
    PPCORE.DLL!03ba5268()   Unknown
    PPCORE.DLL!03ba3358()   Unknown
    PPCORE.DLL!03ba33d5()   Unknown
    PPCORE.DLL!039ec503()   Unknown
    PPCORE.DLL!0325e6bf()   Unknown
    PPCORE.DLL!0325ebe3()   Unknown
    PPCORE.DLL!0325ea1f()   Unknown
    PPCORE.DLL!030758d6()   Unknown
    PPCORE.DLL!033065c8()   Unknown
    PPCORE.DLL!03038b51()   Unknown
    POWERPNT.EXE!00b71567() Unknown
    POWERPNT.EXE!00b714c2() Unknown
    kernel32.dll!763af989() Unknown
    ntdll.dll!779a7084()    Unknown
    ntdll.dll!779a7054()    Unknown
area-Interop-coreclr

Most helpful comment

@jozefizso What a mess. The reason the first path works but the second doesn't has to do with optimizations in the underlying COM system. COM is making an extra QueryInterface call during activation on the "slow" path. The first path is slow and a QueryInterface for the specific interface is performed - in this case IDTExtensibility2. However, on the "fast" path the COM system, correctly mind you, assumes the returned pointer is of the proper type. The proper type for us in .NET Framework was the IClassInterface shenanigans and set up a vtable that handled all the cases.

In .NET Core we removed much of the IClassInterface logic because we didn't support TlbExp scenarios which relied on some incredibly complex logic to fill out and generate the proper IClassInterface in cases when it wasn't defined - normally it wasn't or its purpose would have been in question. We removed all the TlbExp logic when we moved to .NET Core and have no plans on bringing it back.

So where are we? I think I can actually fix this but need to experiment a bunch. We were being too cavalier when we wrote the ComActivator class in .NET 3.x and relied on the built-in marshaller:

https://github.com/dotnet/runtime/blob/c44154b15126786260f8a6397248dd04c0e015e3/src/coreclr/src/System.Private.CoreLib/src/Internal/Runtime/InteropServices/ComActivator.cs#L544-L547

This is a real bug and COM clients like Office can't really do much about it. We will need to figure out the cost of fixing this or a work around. I have some ideas - the first being don't use the built-in marshal by interface mechanism.

/cc @jkoritzinsky @elinor-fung @jeffschwMSFT

All 14 comments

I couldn't figure out the best area label to add to this issue. If you have write-permissions please help me learn by adding exactly one area label.

@jozefizso Thank you for taking the time to file this issue. Would it be possible to supply a trivial add-in that can be used to reproduce the issue? The C# part of the extension would probably be enough.

@AaronRobinsonMSFT yes, I should have sample project, let me find it...

Hi @AaronRobinsonMSFT,
here is the sample project: https://github.com/jozefizso/PowerPointAddin-NetCore3-Crash

It should be compilable and the registration script should work.

@jozefizso Thanks for the project. I am able to reproduce the issue locally. I see what is happening, but am unsure as to the why. The first time the add in is loaded, everything is correct. However on the subsequent load the plugin has a slightly different vtable layout and is corrupting the application's IDispatch pointer - instead of calling OnConnection(), it is calling the object's ToString(). This is 100 % reproducible for me so should be able to pinpoint the issues in the coming days.

Thanks a lot for looking into this.

@jozefizso What a mess. The reason the first path works but the second doesn't has to do with optimizations in the underlying COM system. COM is making an extra QueryInterface call during activation on the "slow" path. The first path is slow and a QueryInterface for the specific interface is performed - in this case IDTExtensibility2. However, on the "fast" path the COM system, correctly mind you, assumes the returned pointer is of the proper type. The proper type for us in .NET Framework was the IClassInterface shenanigans and set up a vtable that handled all the cases.

In .NET Core we removed much of the IClassInterface logic because we didn't support TlbExp scenarios which relied on some incredibly complex logic to fill out and generate the proper IClassInterface in cases when it wasn't defined - normally it wasn't or its purpose would have been in question. We removed all the TlbExp logic when we moved to .NET Core and have no plans on bringing it back.

So where are we? I think I can actually fix this but need to experiment a bunch. We were being too cavalier when we wrote the ComActivator class in .NET 3.x and relied on the built-in marshaller:

https://github.com/dotnet/runtime/blob/c44154b15126786260f8a6397248dd04c0e015e3/src/coreclr/src/System.Private.CoreLib/src/Internal/Runtime/InteropServices/ComActivator.cs#L544-L547

This is a real bug and COM clients like Office can't really do much about it. We will need to figure out the cost of fixing this or a work around. I have some ideas - the first being don't use the built-in marshal by interface mechanism.

/cc @jkoritzinsky @elinor-fung @jeffschwMSFT

In the old days, when using COM Shim Wizard, there was this ManagedAggregator class which had to bind the inner and outer objects for .NET and Shim C++ library to work.

https://github.com/jozefizso/COMShimWizard/blob/master/VCWizards/AddinShim/ManagedAggregator/ManagedAggregator.cs#L45-L55

Is this issue related to that mechanics? Does the .NET Core provide this managed aggregator implementation so we don't have to do it?

@jozefizso Not in this case. There is no aggregation issue with this specific issue, not saying that also doesn't exist but I don't see that solving this problem. We do handle aggregation internally albeit minimally compared to .NET Framework. In .NET Framework the effort made to support the infinite boundary of aggregation is done. In .NET Core we limited that boundary to a somewhat reasonable level - I would be shocked if anyone finds the new limit.

I have a local fix for this. Unsure about .NET 3.1, but I will get it into .NET 5. The .NET 3.1 servicing bar will need to be checked next week. Thanks for taking the time to report the issue.

@jozefizso I have closed this for .NET 5 since that is our latest. The .NET 3.1.x fix will be heading for bar check this week and its progress can be followed at https://github.com/dotnet/coreclr/pull/28073. Again, we really appreciate you filing this issue.

@AaronRobinsonMSFT could it also be a reason of system.data.oledb being unusable https://github.com/dotnet/runtime/issues/36954 ?

@MaceWindu Perhaps. I don鈥檛 know how the OleDb is implemented, but if it involves a managed COM server then it is possible. Note that COM aggregation using a managed object could fall into this path as well. The behavior this issue introduces is unpredictable so knowing for sure is hard. Is there a small repo I can run? A unit test involving XUnit is rather complicated so a small console app is preferable- even if it doesn鈥檛 repo 100 % that is okay.

@AaronRobinsonMSFT thanks a lot for fixing this issue. 馃

@AaronRobinsonMSFT , I'm afraid I don't have simple repro code right now. It fails on big test suite and not on every run, but always in same tests. Running failing tests only doesn't reproduce issue.
Probably I will wait for fix release and retest it.

Was this page helpful?
0 / 5 - 0 ratings