FTL crashed

Created on 27 Feb 2020  路  14Comments  路  Source: pi-hole/FTL

[2020-02-27 09:07:46.522 15352] Resizing "/FTL-dns-cache" from 233472 to 237568
[2020-02-27 09:18:00.508 15352] Resizing "/FTL-clients" from 204800 to 225280
[2020-02-27 09:18:00.524 15352] !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
[2020-02-27 09:18:00.524 15352] ---------------------------->  FTL crashed!  <----------------------------
[2020-02-27 09:18:00.524 15352] !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
[2020-02-27 09:18:00.524 15352] Please report a bug at https://github.com/pi-hole/FTL/issues
[2020-02-27 09:18:00.524 15352] and include in your report already the following details:
[2020-02-27 09:18:00.524 15352] FTL has been running for 160434 seconds
[2020-02-27 09:18:00.524 15352] FTL branch: release/v5.0
[2020-02-27 09:18:00.524 15352] FTL version: vDev-71e8498
[2020-02-27 09:18:00.524 15352] FTL commit: 71e8498
[2020-02-27 09:18:00.524 15352] FTL date: 2020-02-25 08:21:07 +0100
[2020-02-27 09:18:00.525 15352] FTL user: started as pihole, ended as pihole
[2020-02-27 09:18:00.525 15352] Compiled for armhf (compiled on CI) using arm-linux-gnueabihf-gcc (Debian 6.3.0-18) 6.3.0 20170516
[2020-02-27 09:18:00.525 15352] Received signal: Segmentation fault
[2020-02-27 09:18:00.525 15352]      at address: 0xb6dbd5ec
[2020-02-27 09:18:00.525 15352]      with code: SEGV_MAPERR (Address not mapped to object)
[2020-02-27 09:18:00.525 15352] Backtrace:
[2020-02-27 09:18:00.526 15352] B[0000]: 0x474b24, /usr/bin/pihole-FTL(+0x22b24) [0x474b24]
[2020-02-27 09:18:00.526 15352] B[0001]: 0xb6e0c130, /lib/arm-linux-gnueabihf/libc.so.6(__default_rt_sa_restorer+0) [0xb6e0c130]
[2020-02-27 09:18:00.526 15352] B[0002]: 0x479f38, /usr/bin/pihole-FTL(resolveClients+0xb7) [0x479f38]
[2020-02-27 09:18:00.526 15352] B[0003]: 0x47a15e, /usr/bin/pihole-FTL(DNSclient_thread+0x85) [0x47a15e]
[2020-02-27 09:18:00.526 15352] ------ Listing content of directory /dev/shm ------
[2020-02-27 09:18:00.526 15352] File Mode User:Group  Filesize Filename
[2020-02-27 09:18:00.526 15352] rwxrwxrwx root:root 260 .
[2020-02-27 09:18:00.527 15352] rwxr-xr-x root:root 4K ..
[2020-02-27 09:18:00.527 15352] rw------- pihole:pihole 4K FTL-per-client-regex
[2020-02-27 09:18:00.527 15352] rw------- pihole:pihole 238K FTL-dns-cache
[2020-02-27 09:18:00.527 15352] rw------- pihole:pihole 53K FTL-overTime
[2020-02-27 09:18:00.528 15352] rw------- pihole:pihole 5M FTL-queries
[2020-02-27 09:18:00.528 15352] rw------- pihole:pihole 20K FTL-upstreams
[2020-02-27 09:18:00.528 15352] rw------- pihole:pihole 225K FTL-clients
[2020-02-27 09:18:00.529 15352] rw------- pihole:pihole 131K FTL-domains
[2020-02-27 09:18:00.529 15352] rw------- pihole:pihole 201K FTL-strings
[2020-02-27 09:18:00.529 15352] rw------- pihole:pihole 12 FTL-settings
[2020-02-27 09:18:00.529 15352] rw------- pihole:pihole 120 FTL-counters
[2020-02-27 09:18:00.530 15352] rw------- pihole:pihole 28 FTL-lock
[2020-02-27 09:18:00.530 15352] ---------------------------------------------------
[2020-02-27 09:18:00.530 15352] Thank you for helping us to improve our FTL engine!
[2020-02-27 09:18:00.530 15352] FTL terminated!
Bug Bugfix in progress

All 14 comments

Thank you for the log.

If the condition is repeatable, can you follow the guide at https://docs.pi-hole.net/ftldns/debugging/ to get some more detailed information?

Been running in gdb for about 15 hours now, hasn't crashed yet. It's been going down about once a day for the past week, so hopefully we'll catch it soon.

OK, it must have been watching over my shoulder as I typed that last message, because it just segfaulted...

Detaching after fork from child process 24364]
[Detaching after fork from child process 24365]

Thread 1 "pihole-FTL" received signal SIGSEGV, Segmentation fault.
0x004a326c in bindText ()
(gdb) backtrace
#0  0x004a326c in bindText ()
#1  0x004a34de in sqlite3_bind_text ()
#2  0x00422780 in domain_in_list (listname=0x50b318 "whitelist", stmt=<optimized out>, domain=0xb37666f9 "client.wns.windows.com")
    at src/database/gravity-db.c:506
#3  in_whitelist (domain=domain@entry=0xb37666f9 "client.wns.windows.com", client=client@entry=0xb69fb500, clientID=354, clientID@entry=4371287)
    at src/database/gravity-db.c:568
#4  0x0042a4de in _FTL_check_blocking (queryID=queryID@entry=3356978, domainID=domainID@entry=34419376, clientID=4371287, clientID@entry=34948520,
    blockingreason=0x1d0c, blockingreason@entry=0x20d3808, line=<optimized out>, file=<optimized out>) at src/dnsmasq_interface.c:192
#5  0x0042b356 in _FTL_check_blocking (file=0x50e644 "src/dnsmasq_interface.c", line=564, blockingreason=0x20d3808, clientID=34948520,
    domainID=34419376, queryID=<optimized out>) at src/dnsmasq_interface.c:485
#6  _FTL_new_query (flags=<optimized out>, name=<optimized out>, blockingreason=0x20d3808, blockingreason@entry=0xbed16970, addr=<optimized out>,
    types=types@entry=0x21545a8 "query[A]", id=54543, type=type@entry=1 '\001', file=0x5109a8 "src/dnsmasq/forward.c", line=1579)
    at src/dnsmasq_interface.c:564
#7  0x0043bfc4 in receive_query (listen=listen@entry=0x20d8158, now=0, now@entry=1582912214) at src/dnsmasq/forward.c:1578
#8  0x00449f96 in check_dns_listeners (now=now@entry=1582912214) at src/dnsmasq/dnsmasq.c:1658
#9  0x0044b1fc in main_dnsmasq (argc=<optimized out>, argv=<optimized out>) at src/dnsmasq/dnsmasq.c:1108
#10 0x0041eb4a in main (argc=1, argv=<optimized out>) at src/main.c:87

Another one (LMK if I should file separate issues instead of appending):

Thread 253 "telnet-17" received signal SIGSEGV, Segmentation fault.
[Switching to Thread 0xb2aff460 (LWP 3208)]
0xb6f03e44 in __GI___pthread_mutex_lock (mutex=0x45532853) at pthread_mutex_lock.c:67
67      pthread_mutex_lock.c: No such file or directory.
(gdb) backtrace
#0  0xb6f03e44 in __GI___pthread_mutex_lock (mutex=0x45532853) at pthread_mutex_lock.c:67
#1  0x0058a27c in bindText ()
#2  0x0058a4de in sqlite3_bind_text ()
#3  0x00509b82 in domain_in_list (listname=0x5f2338 "auditlist", stmt=0x1931c08, domain=0xb6d6911d "cloudsync-tw.synology.com")
    at src/database/gravity-db.c:506
#4  in_auditlist (domain=0xb6d6911d "cloudsync-tw.synology.com") at src/database/gravity-db.c:589
#5  0x0050bb26 in getTopDomains (client_message=client_message@entry=0xb21005b8 ">top-domains for audit", sock=sock@entry=0xb2afea28)
    at src/api/api.c:303
#6  0x0050ad18 in process_request (client_message=client_message@entry=0xb21005b8 ">top-domains for audit", sock=sock@entry=0xb2afea28)
    at src/api/request.c:51
#7  0x0050a0a2 in telnet_connection_handler_thread (socket_desc=0xb3500708) at src/api/socket.c:336
#8  0xb6f01494 in start_thread (arg=0xb2aff460) at pthread_create.c:486
#9  0xb6e84578 in ?? () at ../sysdeps/unix/sysv/linux/arm/clone.S:73 from /lib/arm-linux-gnueabihf/libc.so.6
Backtrace stopped: previous frame identical to this frame (corrupt stack?)

Not sure what this is:

Thread 1 "pihole-FTL" received signal SIG34, Real-time event 34.
__GI___poll (timeout=-1, nfds=9, fds=0x1513188) at ../sysdeps/unix/sysv/linux/poll.c:29
29      in ../sysdeps/unix/sysv/linux/poll.c
(gdb) backtrace
#0  __GI___poll (timeout=-1, nfds=9, fds=0x1513188) at ../sysdeps/unix/sysv/linux/poll.c:29
#1  __GI___poll (fds=0x1513188, nfds=9, timeout=timeout@entry=-1) at ../sysdeps/unix/sysv/linux/poll.c:26
#2  0x005056b0 in poll (__timeout=__timeout@entry=-1, __nfds=<optimized out>, __fds=<optimized out>)
    at /usr/arm-linux-gnueabihf/include/bits/poll2.h:46
#3  do_poll (timeout=timeout@entry=-1) at src/dnsmasq/poll.c:78
#4  0x005191a2 in main_dnsmasq (argc=<optimized out>, argv=<optimized out>) at src/dnsmasq/dnsmasq.c:1038
#5  0x004ecb4a in main (argc=1, argv=<optimized out>) at src/main.c:87

It seems to happen when I add something to the Blacklist from the Audit Log page (which seems to perform much better than it used to, BTW).

Okay, thanks, I'm traveling this weekend so FTL troubleshooting is difficult, but the crash you're seeing is happening inside sqlite3 subroutines which is, generally, a difficult problem. The crash happens here:
https://github.com/pi-hole/FTL/blob/71e849816b126f14300cde39d6ce4b9b40e651e1/src/database/gravity-db.c#L506
However, your backtrace also tells us that, both, stmt and domain seem to be valid (as in not NULL):

stmt=0x1931c08
domain=0xb6d6911d "cloudsync-tw.synology.com"

There has to be something about stmt going wrong, interestingly enough, this doesn't seem to fit at all to the crash you were reporting initially (first post).

Thread 1 "pihole-FTL" received signal SIG34, Real-time event 34
This is not an issue, you can ignore this signal and simply continue.

Thanks for the update -- LMK if there's anything else I can do. Enjoy the weekend.

I'm having similar problems since today, although I'm on version 4.3.1: The DNS service basically crashes every few seconds. It's virtually unusable at the moment.

Here is an excerpt from my logs:

[2020-03-03 22:52:46.894 17397] Please report a bug at https://github.com/pi-hole/FTL/issues
[2020-03-03 22:52:46.895 17397] and include in your report already the following details:
[2020-03-03 22:52:46.895 17397] FTL has been running for 117 seconds
[2020-03-03 22:52:46.895 17397] FTL branch: master
[2020-03-03 22:52:46.895 17397] FTL version: v4.3.1
[2020-03-03 22:52:46.895 17397] FTL commit: b60d63f
[2020-03-03 22:52:46.895 17397] FTL date: 2019-05-25 21:37:26 +0200
[2020-03-03 22:52:46.895 17397] FTL user: started as pihole, ended as pihole
[2020-03-03 22:52:46.896 17397] Received signal: Segmentation fault
[2020-03-03 22:52:46.896 17397] at address: 0
[2020-03-03 22:52:46.896 17397] with code: SEGV_MAPERR (Address not mapped to object)
[2020-03-03 22:52:46.897 17397] Backtrace:
[2020-03-03 22:52:46.898 17397] Thank you for helping us to improve our FTL engine!
[2020-03-03 22:52:46.898 17397] FTL terminated!

Same issue.

[2020-03-03 16:42:27.351 1198] Please report a bug at https://github.com/pi-hole/FTL/issues
[2020-03-03 16:42:27.351 1198] and include in your report already the following details:
[2020-03-03 16:42:27.351 1198] FTL has been running for 8 seconds
[2020-03-03 16:42:27.351 1198] FTL branch: master
[2020-03-03 16:42:27.351 1198] FTL version: v4.3.1
[2020-03-03 16:42:27.351 1198] FTL commit: b60d63f
[2020-03-03 16:42:27.351 1198] FTL date: 2019-05-25 21:37:26 +0200
[2020-03-03 16:42:27.351 1198] FTL user: started as pihole, ended as pihole
[2020-03-03 16:42:27.351 1198] Received signal: Segmentation fault
[2020-03-03 16:42:27.351 1198] at address: 0
[2020-03-03 16:42:27.351 1198] with code: SEGV_MAPERR (Address not mapped to object)
[2020-03-03 16:42:27.352 1198] Backtrace:
[2020-03-03 16:42:27.352 1198] Thank you for helping us to improve our FTL engine!
[2020-03-03 16:42:27.352 1198] FTL terminated!
[2020-03-03 16:51:45.216 2127] Using log file /var/log/pihole-FTL.log

Can't keep it running long enough to debug with gdb.

Those two are related to #705

Thanks @sylveon you are right!

@awallgren How large is your network roughly? Do you have more than 300 active clients?

If the crash is still happening, please try

pihole checkout ftl tweak/sqlite_debugging

This will not really fix anything, however, it will add some missing details to the backtrace you have obtained above. There's also a small fix in it now, let's see...

I have not seen this recur for quite a while. Feel free to close it.

Was this page helpful?
0 / 5 - 0 ratings

Related issues

JOHRY picture JOHRY  路  17Comments

tronyx picture tronyx  路  17Comments

herdingcatz picture herdingcatz  路  23Comments

DarkDimius picture DarkDimius  路  18Comments

ric-dicle picture ric-dicle  路  26Comments