Runtime: Where to start with Segmentation Fault (core dumped)

Created on 8 Feb 2017  路  26Comments  路  Source: dotnet/runtime

Moving from dotnet/cli#5616 on behalf of @fearthecowboy


Steps to reproduce

I've ported our code base to run on .NET Core, and we've got it working fairly fine on Windows.

I went to run it on Linux (Ubuntu 16.04 ) and I can get it to build, when I go to run it I get Segmentation Fault (core dumped)

Where do I start to figure out whats going wrong?

I added COREHOST_TRACE=1 when I ran it, and it spit out a bunch of information (nothing looked too telling...)

it ends with:

image

Environment data

dotnet --info output:

image

question

Most helpful comment

@fearthecowboy I have just merged in change #9650 that ensures that stack overflow is properly reported in all cases.

All 26 comments

I'd just like to mention that this is a fairly important step as part of our work to ensure that all the Azure SDK and services are built and tested on Linux and MacOS, so if there is someone that can help us out, that'd be really awesome.

this morning, I'll post the steps to build/repro from our code base.

What's the architecture of the machine? First thought based on what I see is that you didn't publish for ubuntu x64.

How are you building the output aND how are you launching it?

I'll respond again shortly...

@fearthecowboy Have you tried running the host under lldb/gdb? They should break at the point where core-dump generation is triggered and you should be able to look at your callstack to determine what is going wrong.

Here's how to replicate what I'm looking at

git clone https://github.com/fearthecowboy/autorest --branch coreclr --single-branch autorest
cd autorest 
dotnet restore AutoRest.sln
dotnet build src/core/AutoRest/AutoRest.csproj
dotnet src/core/AutoRest/bin/Debug/netcoreapp1.0/AutoRest.dll 

Segmentation Fault (core dumped)

@Petermarcu it's Ubuntu x64

image

In the .csproj file, I've set <RuntimeIdentifier>ubuntu.16.04-x64</RuntimeIdentifier>

I've also tried doing

dotnet publish src/core/AutoRest/AutoRest.csproj

and running the published binary:
image

using gdb:

image

backtrace:
image

As far as I can tell, you are doing everything right. @janvorli @gkhanna79 , any advice here?

I originally built it on Windows, targeted ubuntu.16.04-x64 and copied the binaries over to WSL, and it segfaulted. I then installed all tools on WSL, and recompiled, and got the same thing.

Then I finally spun up a brand new VM, installed the tools and built it and got the same thing again.

So, the good news is, that's it's terribly consistent 馃槣

It would appear that it dies in glibc at https://code.woboq.org/userspace/glibc/sysdeps/unix/sysv/linux/getsysstats.c.html#317


306 
307 /* Return the number of pages of total/available physical memory in
308    the system.  This used to be done by parsing /proc/meminfo, but
309    that's unnecessarily expensive (and /proc is not always available).
310    The sysinfo syscall provides the same information, and has been
311    available at least since kernel 2.3.48.  */
312 long int
313 __get_phys_pages (void)
314 {
315   struct sysinfo info;
316 
317   __sysinfo (&info);
318   return sysinfo_mempages (info.totalram, info.mem_unit);
319 }

We're wandering outside my immediate knowledge here...

adding @jkotas in case he has any ideas.

@fearthecowboy let me try to repro it. But before I start, could you please disass the crashing function so that I can see at what instruction it crashed? I wonder if it could be a stack overflow.

Played around, it's dying in different places, depending on which window I run it it.

Not even sure how that's possible.

in this terminal, it's InternalEnterCriticalSection

image

in this terminal it's __get_phys_pages

image

either way, it's commonly parented by WKS::CGHeap::GarbageCollectionGeneration

@Maoni0

@fearthecowboy can you please try to set env variable COMPlus_INTERNAL_ThreadSuspendInjection to 0 and see if the issue goes away?

@janvorli No Effect.

image

Rebuilt it under WSL / ubuntu 14.04 / x64

Still crashes, just somewhere else.

(gdb) r
Starting program: /home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/AutoRest
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".

Program received signal SIGTRAP, Trace/breakpoint trap.
pthread_sigmask (how=2, newmask=<optimized out>, oldmask=0x0) at ../nptl/sysdeps/pthread/pthread_sigmask.c:53
53      ../nptl/sysdeps/pthread/pthread_sigmask.c: No such file or directory.
(gdb) bt
#0  pthread_sigmask (how=2, newmask=<optimized out>, oldmask=0x0) at ../nptl/sysdeps/pthread/pthread_sigmask.c:53
dotnet/core-setup#1  0x00007ffffc62aa24 in lttng_ust_init () from /usr/lib/x86_64-linux-gnu/liblttng-ust.so.0
dotnet/core-setup#2  0x00007fffff41010a in call_init (l=<optimized out>, argc=argc@entry=1, argv=argv@entry=0x7ffffffde0c8, env=env@entry=0x7ffffffde0d8) at dl-init.c:78
dotnet/core-setup#3  0x00007fffff4101f3 in call_init (env=<optimized out>, argv=<optimized out>, argc=<optimized out>, l=<optimized out>) at dl-init.c:36
dotnet/core-setup#4  _dl_init (main_map=main_map@entry=0x6590c0, argc=1, argv=0x7ffffffde0c8, env=0x7ffffffde0d8) at dl-init.c:126
dotnet/core-setup#5  0x00007fffff414c30 in dl_open_worker (a=a@entry=0x7ffffffdc4c8) at dl-open.c:577
dotnet/core-setup#6  0x00007fffff40ffc4 in _dl_catch_error (objname=objname@entry=0x7ffffffdc4b8, errstring=errstring@entry=0x7ffffffdc4c0, mallocedp=mallocedp@entry=0x7ffffffdc4b0,
    operate=operate@entry=0x7fffff414960 <dl_open_worker>, args=args@entry=0x7ffffffdc4c8) at dl-error.c:187
dotnet/core-setup#7  0x00007fffff41437b in _dl_open (file=0x7ffffffdc750 "/home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libcoreclrtraceptprovider.so", mode=-2147483390,
    caller_dlopen=<optimized out>, nsid=-2, argc=1, argv=0x7ffffffde0c8, env=0x7ffffffde0d8) at dl-open.c:661
dotnet/core-setup#8  0x00007fffff1f102b in dlopen_doit (a=a@entry=0x7ffffffdc6e0) at dlopen.c:66
dotnet/core-setup#9  0x00007fffff40ffc4 in _dl_catch_error (objname=0x628c30, errstring=0x628c38, mallocedp=0x628c28, operate=0x7fffff1f0fd0 <dlopen_doit>, args=0x7ffffffdc6e0) at dl-error.c:187
dotnet/core-setup#10 0x00007fffff1f162d in _dlerror_run (operate=operate@entry=0x7fffff1f0fd0 <dlopen_doit>, args=args@entry=0x7ffffffdc6e0) at dlerror.c:163
dotnet/core-setup#11 0x00007fffff1f10c1 in __dlopen (file=<optimized out>, mode=<optimized out>) at dlopen.c:87
dotnet/core-setup#12 0x00007ffffda554bc in PAL_InitializeTracing() () from /home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libcoreclr.so
dotnet/core-setup#13 0x00007fffff41010a in call_init (l=<optimized out>, argc=argc@entry=1, argv=argv@entry=0x7ffffffde0c8, env=env@entry=0x7ffffffde0d8) at dl-init.c:78
dotnet/core-setup#14 0x00007fffff4101f3 in call_init (env=<optimized out>, argv=<optimized out>, argc=<optimized out>, l=<optimized out>) at dl-init.c:36
dotnet/core-setup#15 _dl_init (main_map=main_map@entry=0x6687e0, argc=1, argv=0x7ffffffde0c8, env=0x7ffffffde0d8) at dl-init.c:126
dotnet/core-setup#16 0x00007fffff414c30 in dl_open_worker (a=a@entry=0x7ffffffdcb58) at dl-open.c:577
dotnet/core-setup#17 0x00007fffff40ffc4 in _dl_catch_error (objname=objname@entry=0x7ffffffdcb48, errstring=errstring@entry=0x7ffffffdcb50, mallocedp=mallocedp@entry=0x7ffffffdcb40,
    operate=operate@entry=0x7fffff414960 <dl_open_worker>, args=args@entry=0x7ffffffdcb58) at dl-error.c:187
dotnet/runtime#2412 0x00007fffff41437b in _dl_open (file=0x655a38 "/home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libcoreclr.so", mode=-2147483647, caller_dlopen=<optimized out>,
    nsid=-2, argc=1, argv=0x7ffffffde0c8, env=0x7ffffffde0d8) at dl-open.c:661
dotnet/core-setup#19 0x00007fffff1f102b in dlopen_doit (a=a@entry=0x7ffffffdcd70) at dlopen.c:66
dotnet/runtime#2413 0x00007fffff40ffc4 in _dl_catch_error (objname=0x628c30, errstring=0x628c38, mallocedp=0x628c28, operate=0x7fffff1f0fd0 <dlopen_doit>, args=0x7ffffffdcd70) at dl-error.c:187
dotnet/core-setup#21 0x00007fffff1f162d in _dlerror_run (operate=operate@entry=0x7fffff1f0fd0 <dlopen_doit>, args=args@entry=0x7ffffffdcd70) at dlerror.c:163
dotnet/core-setup#22 0x00007fffff1f10c1 in __dlopen (file=<optimized out>, mode=<optimized out>) at dlopen.c:87
dotnet/runtime#2414 0x00007ffffdeda130 in pal::load_library(char const*, void**) () from /home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libhostpolicy.so
dotnet/core-setup#24 0x00007ffffdec67ef in coreclr::bind(std::string const&) () from /home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libhostpolicy.so
dotnet/runtime#2415 0x00007ffffdebb6c4 in run(arguments_t const&) () from /home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libhostpolicy.so
dotnet/runtime#2416 0x00007ffffdebc592 in corehost_main () from /home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libhostpolicy.so
dotnet/runtime#2417 0x00007ffffe18154f in execute_app(std::string const&, corehost_init_t*, int, char const**) () from /home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libhostfxr.so
dotnet/core-setup#28 0x00007ffffe188255 in fx_muxer_t::read_config_and_execute(std::string const&, std::string const&, std::unordered_map<std::string, std::vector<std::string, std::allocator<std::string> >, std::hash<std::str
ing>, std::equal_to<std::string>, std::allocator<std::pair<std::string const, std::vector<std::string, std::allocator<std::string> > > > > const&, int, char const**, host_mode_t) ()
   from /home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libhostfxr.so
dotnet/core-setup#29 0x00007ffffe187601 in fx_muxer_t::parse_args_and_execute(std::string const&, std::string const&, int, int, char const**, bool, host_mode_t, bool*) ()
   from /home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libhostfxr.so
dotnet/core-setup#30 0x00007ffffe1886e9 in fx_muxer_t::execute(int, char const**) () from /home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libhostfxr.so
dotnet/core-setup#31 0x00007ffffe1815c5 in hostfxr_main () from /home/garretts/autorest/src/core/AutoRest/bin/Debug/netcoreapp1.0/ubuntu.14.04-x64/publish/libhostfxr.so
dotnet/core-setup#32 0x000000000040b6db in run(int, char const**) ()
dotnet/runtime#2418 0x000000000040b810 in main ()
(gdb)

@fearthecowboy I can confirm that I can repro the issue on my Ubuntu 14.04 as well.

It is a stack overflow. That's why you hit it at different places. It seems that your stack trace above is unrelated to the issue.
In my stack trace, I could see that the code was trying to make a call, but the RSP was already out of stack:
At frame 0:

0x7ffff62c4fa2 <+1490>: callq  0x7ffff603e320

(lldb) x/gx $rsp
error: memory read failed for 0x7fffff7fee00

Breaking into the debugger right after starting the app and dumping the managed stack, I can see that there is an infinite recursion in AutoRest.Core.Utilities.FileSystem.ReadFileAsText(System.String) calling itself.

Part of the stack to demonstrate that:

00007FFFFFFFB950 00007FFF7CEA10BF AutoRest.Core.dll!AutoRest.Core.Utilities.FileSystem.ReadFileAsText(System.String) + 223
00007FFFFFFFBA30 00007FFF7CEA10BF AutoRest.Core.dll!AutoRest.Core.Utilities.FileSystem.ReadFileAsText(System.String) + 223
00007FFFFFFFBB10 00007FFF7CEA10BF AutoRest.Core.dll!AutoRest.Core.Utilities.FileSystem.ReadFileAsText(System.String) + 223
00007FFFFFFFBBF0 00007FFF7CEA10BF AutoRest.Core.dll!AutoRest.Core.Utilities.FileSystem.ReadFileAsText(System.String) + 223
00007FFFFFFFBCD0 00007FFF7CEA10BF AutoRest.Core.dll!AutoRest.Core.Utilities.FileSystem.ReadFileAsText(System.String) + 223
00007FFFFFFFBDB0 00007FFF7CEA10BF AutoRest.Core.dll!AutoRest.Core.Utilities.FileSystem.ReadFileAsText(System.String) + 223
00007FFFFFFFBE90 00007FFF7CEA10BF AutoRest.Core.dll!AutoRest.Core.Utilities.FileSystem.ReadFileAsText(System.String) + 223
00007FFFFFFFBF70 00007FFF7CE9EE9D AutoRest.Core.dll!AutoRest.Core.Extensibility.ExtensionsLoader.GetConfigurationFileContent(AutoRest.Core.Settings) + 797
00007FFFFFFFC060 00007FFF7CE81BE9 AutoRest.dll!AutoRest.HelpGenerator.Generate(System.String, AutoRest.Core.Settings) + 2377
00007FFFFFFFC6E0 00007FFF7CE718DB AutoRest.dll!AutoRest.Program.Main(System.String[]) + 1947

It's actually getting into our code?!

Well, that's good to know.

ok, lemme go look at that...

Thanks!, that unblocked me!

So, I'm a little confused how someone is supposed to know how to find out what's going on in a case like this.

In a 'traditional' dotnet app, I'd generally expect a StackOverflowException to get thrown, and give me a smidgen of a chance of finding out where the actual error is.

I guess that you could use VSCode to debug the managed code on Unix for better experience.

The issue with stack overflow in a tight loop is that when we get out of stack, the process doesn't have enough stack to even execute the SIGSEGV handler. So the process is killed without any notification that would allow us to print some message.
There is a way around it, but it introduces some ugly complexity into handling of other issues that manifest itself as SIGSEGV, like e.g. a null reference exception. I have an issue https://github.com/dotnet/coreclr/issues/1504 assigned to myself to see if we could reasonably do that.

@janvorli, Would this same stack overflow in user code happen on Windows? From a .NET developers perspective, getting a segfault for a problem in their code is the worst possible outcome. Do we need to think about things like reserving stack space to be able to not segfault and instead give a stack overflow error?

It would not happen on Windows, since the SO handling works in a different way. As for reserving stack space, just imagine a tight loop of a recursive function. The stack grows by the size of the stack frame of that function in small increments until the processor hits the end of the stack and cannot write to the memory. Since the kernel would need to run the SIGSEGV handler on the same stack and there is no space left, the kernel just kills the process and SIGSEGV is reported.
Please note that JIT emits stack probes for frames that are larger than 4kB. If such probe finds the frame would not fit the stack, we display a message, since we have enough stack for doing that.
As I've said, there is a way around it. I have described it in the https://github.com/dotnet/coreclr/issues/1504. We have discussed it in the past and the general feeling was that there is only a little difference between displaying a message and aborting the process and letting the process die with SIGSEGV and it is probably not worth the complexity introduced by the workaround. But I have still left the issue open to eventually give it a try at some point.
Since this is the third time in the last few weeks when someone has asked about it, I guess it is time to bite the bullet and finally give it a try.

I'm unblocked, so you can close this unless you want to track it for frustration purposes :D

image

I'm satisfied with my care!

@fearthecowboy I have just merged in change #9650 that ensures that stack overflow is properly reported in all cases.

there is conflicting libssl packages just do

sudo apt remove libssl1.0.0/now
that should fix it

Was this page helpful?
0 / 5 - 0 ratings