• Services shutdown waits indefinitely on clients blocked in Socket.recv

    From Rob Swindell@1:103/705 to GitLab issue in main/sbbs on Thu Sep 24 18:35:15 2026
    open https://gitlab.synchro.net/main/sbbs/-/work_items/1255

    ## Summary

    On shutdown, the services thread waits indefinitely for dynamic service
    clients to disconnect, and a client blocked in `Socket.recvline()` does not notice the shutdown until its read timeout expires. NNTP waits up to 300 seconds, so an idle NNTP session holds shutdown for up to five minutes, longer than systemd's default 90-second stop timeout. The process is then killed with SIGKILL.

    While it waits, `sbbscon` logs only `Services thread still running`, without naming the services or clients involved.

    ## Observed (2026-09-24, sbbs on Linux)

    ```
    15:48:17 srvc 0000 Services terminate
    15:48:19 srvc 0000 Closing service sockets
    15:48:19 srvc 0000 Waiting for 3 clients to disconnect
    15:48:19 srvc 0004 Gopher !JavaScript warning /sbbs/exec/load/cga_defs.js line 90: Terminated
    15:48:27 Services thread still running
    ... every 10 seconds ...
    15:49:47 Services thread still running
    15:49:47 systemd: sbbs.service: State 'stop-sigterm' timed out. Killing.
    ```

    `Done waiting for clients to disconnect` never appeared. Every other server had terminated by 15:48:21.

    Two of the likely three clients were NNTP sessions from 83.238.214.180
    (sockets 0077 and 0086). Both issued their last command at 15:47:19, and at 15:47:30 both logged `Closing open message base ... due to client inactivity`. From then on `nntpservice.js` reads the next command with `client.socket.recvline(1024, 300)`: 300 seconds, because no message base is open. Neither session logged anything after 15:47:30. The earliest those reads could time out is about 15:52:30, well after systemd's SIGKILL at 15:49:47.

    The third client is not identified. Gopher runs at `LogLevel=Notice`, so its connect and thread-exit lines are suppressed, and whether socket 0004 exited after its `Terminated` warning cannot be read from the log.

    ## Why the read doesn't end

    `services_terminate()` sets `terminated` and each `service[].terminated`. Shutdown then closes the *listening* sockets and polls `active_clients()` every 500 ms, with no time limit (`services.cpp`, "Wait for Dynamic Service
    Threads to terminate"). Client sockets are left open.

    A script notices shutdown only in `js_OperationCallback()`, which runs while JavaScript is executing. `js_recvline()` (`js_socket.cpp`) suspends the request and loops over `js_sock_read_check()` in 1-second `socket_check()` steps until data arrives, the peer disconnects, or the caller's timeout expires. It never looks at the termination flag, so the script does not run again until the read returns. `Socket.recv()`, `poll()` and the other blocking socket methods probably behave the same way.

    ## Suggested direction

    1. Let blocking socket reads in service scripts end early on shutdown. For
    example, `js_sock_read_check()` could check the script's
    `callback.terminated` flag (already wired to `service->terminated` for
    dynamic clients) and return as it does for a disconnect. Alternatively,
    `shutdown()` the client sockets after the listening sockets are closed.
    2. Consider bounding the wait for clients, so one stuck client can't hold
    shutdown past the service manager's stop timeout. (The FTP server's "Waiting
    for N child threads to terminate" is unbounded too; its sessions exited
    within seconds here.)
    3. Say who is being waited on. While waiting, log each service with a non-zero
    client count (`service[i].protocol`, `service[i].clients`), and in `sbbscon`
    replace the bare `Services thread still running` with something that
    identifies what is still running, as the FTP, Web and Mail lines at least
    report their inactivity timeouts.

    ## Not established

    No backtrace was taken during the hang, so the recvline stall is inferred from the log timeline and the code, not observed in a stack. The same shutdown can be reproduced by connecting to NNTP, going idle for more than 10 seconds, and stopping the service; `gdb -p <pid> -batch -ex "thread apply all bt"` during the wait would confirm it.

    -- *Authored by Claude (Claude Code), on behalf of @rswindell*
    --- SBBSecho 3.37-Linux
    * Origin: Vertrauen - [vert/cvs/bbs].synchro.net (1:103/705)