Runtime: Remove the pid argument from createdump command line and only take dumps on the parent process

Created on 23 Sep 2020  路  11Comments  路  Source: dotnet/runtime

Change createdump across all our platforms to only take dumps of its parent process. Display a warning message if a pid is passed on the command line.

Directly using createdump for adhoc dumps is not needed anymore because dotnet-dump collect is available. In both unhandled exception triggered and dotnet-dump collect the runtime launches createdump with the runtime process as the parent.

area-Diagnostics-coreclr

Most helpful comment

Directly using createdump for adhoc dumps is not needed anymore because dotnet-dump collect is available.

createdump has still a huge advantage over dotnet-dump: it does not require the SDK to be downloaded on the target machine.

All 11 comments

Tagging subscribers to this area: @tommcdon
See info in area-owners.md if you want to be subscribed.

Directly using createdump for adhoc dumps is not needed anymore because dotnet-dump collect is available.

createdump has still a huge advantage over dotnet-dump: it does not require the SDK to be downloaded on the target machine.

As @kevingosse said, createdump ships with the runtime so no need to install anything else (especially requiring downloading and installing the whole .NET SDK on a production machine or a small container does not make any sense).
Behind the scene, dotnet-dump is sending a command through EventPipe to the CLR to call createdump: this is why restricting to the parent process would work. But this is just an implementation detail that brings no value.
Shipping SOS outside of the CLR was a mistake in term of deployment and ease of use: don't repeat the same error with createdump ;^)

/cc: @noahfalk, @tommcdon

Hello!

How would this work if the parent process is completely broken, and could not spawn a create-dump child ?

I had this kind of issues in my experience, f.e. all threads stuck due to a bug in unmanaged library, or CPU usage is 100% and the app doesn't respond. In this case the diagnostic tool is my last resort, and I expect it to be as reliable as possible (i.e. not rely on the app process to answer through EventPipes).

Hello!

Could you share what the goal is and what the benefits of this change would be so we can balance them with the foreseen drawbacks that we see in the comments already ?

a mistake in term of deployment and ease of use...

Could you share what the goal is and what the benefits of this change would be...

The security team at Microsoft does regular exploration and they were worried this tool has attack surface larger than it needs to be which is creating unneeded risk. This is a move to mitigate that concern. Tools that are deployed separately (like dotnet-dump) allow .NET users to decide if they want to opt-in to larger attack surfaces that installing additional software in production creates. Tools that ship in-box need to bias more strongly to security risk mitigation even if that means sacrificing diagnostic ease of use. If there is more we could do to lessen the diagnostic part of the impact we are glad to talk about it, but we aren't going to be able to leave createdump as-is.

Also totally get that folks wouldn't want to install the entire SDK to install a dump collection tool. We've got two devs working on https://github.com/dotnet/diagnostics/issues/344 right now and I am fully expecting that to be in a better place before these createdump changes happen. I expect easier access to dotnet-dump will take some of the pressure off of createdump and make this change less painful. Hope that helps!

I suppose this is something to be discussed with the security team, but I have a hard time understanding how createdump could possibly increase the surface attack on Linux. All it does is reading from /proc/pid/mem and executing a few ptrace calls, all of which could be done using the default Linux command-line tools, with the exact same privileges.

The area of attack is in the CLR itself (dotnet-dump is just sending the right command to the other process CLR through EventPipe - https://github.com/dotnet/diagnostics/blob/master/documentation/design-docs/ipc-protocol.md) so it is too risky to install the CLR (just kidding ;^). I don't think that sending this command requires more privileges than dumping the memory of the other process as @kevingosse mentioned.

We have the same Security team at the office so I understand how complicated it is to mitigate this kind of request... For example, we had to add authentication to our equivalent to HTTP-based dotnet-dump.

tools that ship in-box need to bias more strongly to security risk mitigation even if that means sacrificing diagnostic ease of use. If there is more we could do to lessen the diagnostic part of the impact we are glad to talk about it, but we aren't going to be able to leave createdump as-is.

You are not talking about "sacrificing diagnostic ease of use" or "lessen the diagnostic part" but just removing the feature. Having to download another MS tool (here dotnet-dump) to do the exact same thing is not reducing the area of attack because dev/administrators will have to do it anyway and so it will be available for the attacker as createdump is today.

is not reducing the area of attack because dev/administrators will have to do it anyway

I don't think this is accurate. Allowing users to decide if and when to deploy a tool gives at least two benefits:

  • Users that don't use the tool won't deploy it. Risk is now limited only to the subset of developers who do want to use the tool.
  • Users who do use a tool still might decide to only deploy it selectively. For example they might only deploy it on-demand in response to a specific issue, they might only to deploy it to certain machines, and they might use procedural hurdles such as authentication, auditing, or disabling machines from the production traffic before the deployment happens. All of those mechanisms mitigate risk.

You are not talking about "sacrificing diagnostic ease of use" or "lessen the diagnostic part" but just removing the feature

I think the phrasing all depends on how you scope it. If you define the problem narrowly "I want to invoke createdump at the command-line to take a dump of process X" then yeah, that is no longer possible. What I am hoping is that most customers will define the workflow more broadly "I want to invoke _something_ at the command-line to take a dump of process X". In other words it is accomplishing the task which is important, not the exact tool that is used to do it. 'sacrificing ease of use" is because previously there was no additional tool installation required and with dotnet-dump now there would be. That extra step has made the workflow a bit harder, but it shouldn't completely prevent it from working.

Then again, createdump is just doing some things the user could do with other tools readily available on Linux and with the same privileges. So it's just a shortcut, a helper. Any attacker that will be in a position to "exploit" createdump can script those things easily. So removing it effectively removes no risk.

Was this page helpful?
0 / 5 - 0 ratings