On Unix-like platforms (macOS, Linux) macOS only, where (I presume) the exec() family of system calls underlies System.Diagnostics.Process.Start() , certain error conditions that result in the _fundamental inability_ to create a process for the specified executable file are currently quietly masked:
I'm using PowerShell Core for simpler reproduction (macOS only - on Linux, the behavior is as expected); in fact, we discovered the issue there:
# Trying to invoke a nonexistent executable throws an exception,
# as expected.
[System.Diagnostics.Process]::Start('/no/such/executable')
# By contrast, trying to invoke an *invalid* executable returns a process
# object, even though the process for the *target executable* wasn't created.
# (The example creates a plain-text file that is marked as executable, but cannot be
# executed, because it lacks interpreter information (no shebang line)).
'an executable without shebang line' > ($tmpExe = "/tmp/$PID")
chmod +x $tmpExe
[System.Diagnostics.Process]::Start($tmpExe)
Remove-Item $tmpExe
The above yields:
Exception calling "Start" with "1" argument(s): "No such file or directory"
# ...
NPM(K) PM(M) WS(M) CPU(s) Id SI ProcessName
------ ----- ----- ------ -- -- -----------
0 0.00 0.00 0.00 15373 0
The fundamental failure to create the target process due to _nonexistence_ threw an exception, as expected.
The fundamental failure to target process due to _an invalid executable_ was quietly ignored, and a (useless) process object was returned.
8 (ENOEXEC), as the target-process-that-was-never-created's _exit code_.Other error conditions indicated by the exec family of functions - as standardized by POSIX here - may cause the same unexpected behavior.
Notably, it also happens if a shebang line exists, but points to a non-existent interpreter (in which the _system call's errorno_ is 2 (ENOENT)), which is again misreported as the ostensibly successfully created process' exit code.
Thanks for labeling, @joshfree; note that the bug affects os-mac-os-x too.
While it's possible there's a better solution, this behavior was implemented on purpose. By the time exec is Invoked, it's already in the forked child process running separately from the parent: the Process object represents that child. If we're unable to communicate the error back to the parent as part of the fork/exec process, we treat the error as the process's exit code.
@stephentoub:
it's already in the child process running separately from the parent
Yes, but that is an _implementation detail_.
This child is useless from the perspective of the caller and the intent of the .Start() method:
What matters with regard to the _intent_ of the call is is whether a child process was able to be created for the _target executable_.
Reporting the platform-incidental child process that is unrelated to the target executable amounts to a leaky abstraction.
Given the above, can you explain:
what is on purpose about the current behavior?
if there are concerns about changing the behavior (why would anyone rely on the current behavior)?
@stephentoub @mklement0 did we ever get an answer to:
Given the above, can you explain:
- what is on purpose about the current behavior?
- if there are concerns about changing the behavior (why would anyone rely on the current behavior)?
Thanks for following up, @wtgodbe - I guess @stephentoub will need to answer that.
Since you're concerned with "implementation details", what implementation do you propose to address your concerns?
Since you're concerned with "implementation details",
On the contrary: My concern is that the details of the existing implementation are affecting the behavior, creating a leaky abstraction:
As a user, I want to know whether a process I tried to create was created successfully.
If the API didn't actually create the process _with the intended executable_, it is unhelpful and misleading to report a behind-the-scenes _helper_ process instead (that is what I meant by implementation detail: that new processes are created by first forking the calling process and then calling an exec*() system function to replace the forked process' image is how things happen to work on Unix platforms).
If you don't think that's a problem, there is no need to start a conversation about how to implement a possible fix.
If you know that the behavior is _not fixable for technical reasons_ (I can't tell), do tell us, but note that this is quite different from functionality that was "implemented on purpose".
Also, the unexpected behavior then deserves documenting.
If you know that the behavior is not fixable for technical reasons (I can't tell), do tell us, but note that this is quite different from functionality that was "implemented on purpose".
I'm not familiar with a solution to achieve your desired semantics, in particular on macOS where pipe2 isn't available. And the behavior you see was "implemented on purpose" because of how the OS works with regards to fork.
because of how the OS works with regards to fork.
Understood, but that's the aforementioned platform-specific implementation detail that shouldn't affect the behavior of the API.
I'm not familiar with a solution to achieve your desired semantics, in particular on macOS
It seems that bash has found a solution, even on macOS, because it already does detect the system's inability to invoke a shebang-less executable plain-text script and falls back to interpreting the file _itself_; from a comment on https://stackoverflow.com/a/24944521/45375:
Bash interprets scripts without a
#!itself because it callsexecve, and checks if it failed because of a missing#!, executing it itself if so (source).
because it already does detect the system's inability to invoke a shebang-less executable plain-text script and falls back to interpreting the file itself;
There's a difference between a) doing a check on the target file ahead of time and b) fork/exec'ing and special-case handling a resulting error code. I've been talking about the difficulties of doing (b), because that's what this issue focused on. If you know of a good way to address (b), please feel free to submit a PR. If you're talking about (a), which is what it sounds like to me from that stackoverflow post, then it'd be helpful for you to explain what kinds of checks you're interested in seeing added prior to the fork; we would collectively need to evaluate whether those checks provide more benefit than harm.
(a) wouldn't be appropriate, and it's not even what Bash does - it tries execve() _first_ and _then_ falls back to interpreting the file itself (if it makes sense).
Either way, (b) is all we need here and, as it turns out, despite what I initially claimed, (b) does already work on _Linux_ - _it is only macOS that has the problem_ - sorry for not checking that up front (I had just assumed); I've updated this issue's title and the initial post accordingly; please remove the os-linux label.
As for how Bash manages to do (b) even on macOS:
I accidentally posted the wrong Bash source-code link before (since corrected), but here it is again to be safe: http://git.savannah.gnu.org/cgit/bash.git/tree/execute_cmd.c?id=35751626#n5183
Here's the relevant snippet:
SETOSTYPE (0); /* Some systems use for USG/POSIX semantics */
execve (command, args, env);
i = errno; /* error from execve() */
CHECK_TERMSIG;
SETOSTYPE (1);
/* If we get to this point, then start checking out the file.
Maybe it is something we can hack ourselves. */
In other words, a failed execve() call _that returns_ implies failure, with the specific error condition reported in the global errno variable, which Bash caches in variable i and subsequently analyzes and takes action on.
As an aside:
While falling back to a default action - such as interpretation by Bash itself - is clearly not appropriate for System.Diagnostics.Process.Start(), Bash also helpfully transforms the reported execve() errors into _more meaningful errors_ by doing its own sleuthing - something that .NET Core could also benefit from - but that's definitely a nice-to-have.
For example, on Linux, where execve() failures _are_ properly reported as exceptions, the following two distinct error conditions currently result in the _same_ exception, System.ComponentModel.Win32Exception (2): No such file or directory:
The target executable doesn't exist.
The target executable exists, but its shebang line points to a non-existent interpreter.
Bash transforms the latter error into the following, more specific and helpful error message:
/foo/bar: bad interpreter: No such file or directory
Re submitting a PR:
Time permitting, I can take a stab at it, but it'll take me a while to get set up and learn the ropes.
Therefore, my preference is for someone else to take this on - it sounds like it wouldn't be too hard for someone already familiar with the code base.
does already work on Linux - it is only macOS that has the problem
This is why I commented on pipe2. If pipe2 is available, we create a pipe that'll be closed when the child process execs:
https://github.com/dotnet/corefx/blob/443723f8f16d25160f6f0bec3ef1505844fbdf46/src/Native/Unix/System.Native/pal_process.c#L212-L214
and if the exec fails, we write the error value to that pipe:
https://github.com/dotnet/corefx/blob/443723f8f16d25160f6f0bec3ef1505844fbdf46/src/Native/Unix/System.Native/pal_process.c#L255
The parent process waits on the pipe in order to know when the child has exec'd, and if it's able to read an error value from the pipe, uses that as the error code:
https://github.com/dotnet/corefx/blob/443723f8f16d25160f6f0bec3ef1505844fbdf46/src/Native/Unix/System.Native/pal_process.c#L282
This is how the child communicates the error failure to the parent. pipe2 is available on Linux, and things work out nicely: we can communicate the error value to the parent process as part of Start, we can avoid returning from Start until the process has transitioned to be the target process such that it'll have the right Name, etc.
On macOS, pipe2 doesn't exist. As such, we lack a good way of communicating back from the child to the parent in a robust fashion, and as the comment here indicates, we can't just simulate it with fcntl:
https://github.com/dotnet/corefx/blob/443723f8f16d25160f6f0bec3ef1505844fbdf46/src/Native/Unix/System.Native/pal_process.c#L207-L211
In other words, a failed execve() call that returns implies failure, with the specific error condition reported in the global errno variable, which Bash caches in variable i and subsequently analyzes and takes action on.
As I've noted, the call to exec in our case is in the child process, not the same process calling Start. I don't see how the comments here about Bash are relevant. How do you propose to block the parent process until the child process has either successfully exec'd or until the exec fails, and how do you propose to propagate the error value from exec in the case of failure back to the parent?
Thanks for the detailed background info, @stephentoub.
How do you propose to block the parent process until the child process has either successfully exec'd or until the exec fails, and how do you propose to propagate the error value from exec in the case of failure back to the parent?
_Personally_, I have no idea - you clearly know much more about this than I do.
I naively assumed that the Bash solution would be portable, but I hadn't considered the cross-process sync issues.
However, this takes us back to my earlier question, which _I_ cannot answer:
Is this not fixable _for technical reasons_?
If so, let's document the problem and call it a day.
I do hope it is clear by now that it _is_ a problem, and perhaps a less exotic scenario in which it occurs is if an executable shell script's shebang line happens to point to a _non-existent interpreter_ (I've also added that to the initial post).
I can try to look into this further to see if there's a solution, but it sounds like I'd be playing catch-up.
It sounds like you, @stephentoub, aren't aware of a solution, but are there other subject-matter experts we can ping to weigh in? Or are you confident enough to make the call that there simply is no solution?
Is this not fixable for technical reasons?
It was implemented as it is implemented on macOS because on macOS we do not currently have a reliable and efficient way to communicate the information back to the parent process. Ideally, we would mimic the behavior of the code as it's historically worked on Windows, where any failure results in an exception from Start. But when we can't make that happen, we fall back to still communicating the failure as the Process' exit code, since that is in fact what's happening.
If someone comes up with a scheme that allows us to improve this, we'll happily consider the PR. In the meantime, though, it's behaving as intended.
are there other subject-matter experts we can ping to weigh in?
@tmds
As @stephentoub explained, we rely on CLOEXEC to be notified of the exec and mac doesn't provide an atomic way to create a fd that has the flag set.
Maybe we could create the CLOEXEC pipe in the child process. We can use another pipe created by the parent to pass a pipe end from the child to the parent.
int childToParent[2] = pipe();
childToParent[] set CLOEXEC // best-effort
fork();
if (child)
{
int waitForChildToExecPipe[2] = pipe();
waitForChildToExecPipe[writeEnd] set CLOEXEC
sendfd(childToParent[readEnd], waitForChildToExecPipe[readEnd]);
close(waitForChildToExecPipe[readEnd]);
...
execve();
write(waitForChildToExecPipe[writeEnd], errno);
exit();
}
else // parent
{
int waitForChildToExecFd = recvfd(childToParent[readEnd]);
...
}
Thanks for the suggestion, @tmds.
Forgive me, if the following question has an obvious answer or betrays a lack of understanding, I'm out of my depth here.
Wouldn't the lack of atomicity of setting CLOEXEC at pipe-creation time still be a problem (the pipe fds potentially leaking to concurrently forked processes before fcntl can be called with FD_CLOEXEC), now in 2 places?
Generally, when you say "Maybe we could" do you mean to say you're not sure it will work robustly?
Or are there performance or other concerns?
Those are good questions.
Wouldn't the lack of atomicity of setting CLOEXEC at pipe-creation time still be a problem
Since Mac doesn't have apis to atomically set CLOEXEC when creating new file descriptors, some leaking is already possible. It's not something we can solve.
For this particular case, we don't want to leak the pipe-end we use for writing back the error code to the parent into another child. If that happened then we'd only detect the close when that other child has terminated. By creating it in the child process, we have a write end that can't be inherited by other child processes.
Generally, when you say "Maybe we could" do you mean to say you're not sure it will work robustly?
Or are there performance or other concerns?
I think it should work, but maybe I'm missing something... the devil is in the detail. Or maybe some things don't work as expected on Mac.
I appreciate the explanation, @tmds.
Would you be willing to give this a try? (As stated, I'm not the best candidate.)
Most helpful comment
@stephentoub:
Yes, but that is an _implementation detail_.
This child is useless from the perspective of the caller and the intent of the
.Start()method:What matters with regard to the _intent_ of the call is is whether a child process was able to be created for the _target executable_.
Reporting the platform-incidental child process that is unrelated to the target executable amounts to a leaky abstraction.
Given the above, can you explain:
what is on purpose about the current behavior?
if there are concerns about changing the behavior (why would anyone rely on the current behavior)?