Package: x2goclient; Maintainer for x2goclient is X2Go Developers <x2go-dev@lists.x2go.org>; Source for x2goclient is src:x2goclient.
Reported by: Philipp Klenze <p.klenze@gsi.de>
Date: Thu, 17 Sep 2026 08:50:02 UTC
Severity: normal
Found in version 4.1.2.3-4
Reply or subscribe to this bug.
View this report as an mbox folder, status mbox, maintainer mbox
Report forwarded
to x2go-dev@lists.x2go.org, X2Go Developers <x2go-dev@lists.x2go.org>:
Bug#1637; Package x2goclient.
(Thu, 17 Sep 2026 08:50:02 GMT) (full text, mbox, link).
Acknowledgement sent
to Philipp Klenze <p.klenze@gsi.de>:
New Bug report received and forwarded. Copy sent to X2Go Developers <x2go-dev@lists.x2go.org>.
(Thu, 17 Sep 2026 08:50:02 GMT) (full text, mbox, link).
Message #5 received at submit@bugs.x2go.org (full text, mbox, reply):
[Message part 1 (text/plain, inline)]
Package: x2goclient
Version: 4.1.2.3-4
--------------------------------------------------------------------
NOTE ON HOW THIS REPORT WAS PRODUCED
--------------------------------------------------------------------
The debugging (gdb sessions, valgrind run, log analysis) was done
together with an LLM assistant (Claude) at my direction -- I ran
every command myself and verified the findings, but the writeup
itself is LLM-coauthored. Flagging this in the subject line;
disregard or downgrade this note as you see fit.
--------------------------------------------------------------------
SUMMARY
--------------------------------------------------------------------
x2goclient intermittently segfaults with a SIGSEGV inside libssh's
ssh_select() -> ssh_event_add_session(), called from
SshMasterConnection::channelLoop(). Both a gdb post-mortem analysis
of a core file and a live valgrind/memcheck run independently point
to the same thing: the channels[] array that channelLoop() hands to
ssh_select() gets corrupted, almost certainly by a race between the
SSH worker thread iterating it and something else (channel
teardown?) mutating/freeing it concurrently.
--------------------------------------------------------------------
ENVIRONMENT
--------------------------------------------------------------------
- x2goclient: 4.1.2.3-4 (Debian trixie)
- libssh4: 0.11.2-1+deb13u1 (trixie-security), reported SONAME
libssh.so.4.10.5
- OS: Debian trixie, amd64
- Connection topology: two-hop, via an SSH jump host
(SshMasterConnection instantiated once for the direct hop to the
proxy, again for the tunnelled hop to the target host)
- valgrind: 3.24.0
--------------------------------------------------------------------
BACKTRACE (gdb, from a core file)
--------------------------------------------------------------------
Program terminated with signal SIGSEGV, Segmentation fault.
#0 ssh_select (channels=0x7fc590005df0, outchannels=0x7fc590009dc0,
maxfd=0,
readfds=0x7fc597349910, timeout=<optimized out>) at
./src/connect.c:351
#1 0x0000564b01ded639 in SshMasterConnection::channelLoop (
this=0x7fc5c8007740) at ../src/sshmasterconnection.cpp:1981
#2 0x0000564b01df1304 in SshMasterConnection::run (this=0x7fc5c8007740)
at ../src/sshmasterconnection.cpp:752
#3 0x00007fc5d32ded21 in operator() (__closure=<optimized out>)
at thread/qthread_unix.cpp:350
#4 (anonymous
namespace)::terminate_on_exception<QThreadPrivate::start(void*)::<lambda()>
> (t=...) at thread/qthread_unix.cpp:287
#5 QThreadPrivate::start (arg=0x7fc5c8007740) at
thread/qthread_unix.cpp:310
#6 0x00007fc5d2c9e8bb in start_thread (arg=<optimized out>)
at ./nptl/pthread_create.c:448
#7 0x00007fc5d2d1c538 in __GI___clone3 ()
at ../sysdeps/unix/sysv/linux/x86_64/clone3.S:78
libssh's connect.c at that point (0.11.2 source):
346 ssh_event event = ssh_event_new();
347 int firstround = 1;
348
349 base_tm = tm = (timeout->tv_sec * 1000) +
(timeout->tv_usec / 1000);
350 for (i = 0 ; channels[i] != NULL; ++i) {
351 ssh_event_add_session(event, channels[i]->session);
352 }
The crash is on line 351, dereferencing channels[i]->session.
(I reconstructed this snippet from general knowledge of the libssh
0.11.x source rather than fetching the exact trixie-security tarball
-- worth double-checking against the real source for this build if
it matters to you.)
--------------------------------------------------------------------
POINTER INSPECTION
--------------------------------------------------------------------
(gdb) print channels[0]
$9 = (ssh_channel) 0x7fc590005
(gdb) print channels[1]
$10 = (ssh_channel) 0x0
(gdb) print *channels[0]
Cannot access memory at address 0x7fc590005
For reference, the array itself lives at channels = 0x7fc590005df0.
channels[0]'s value (0x7fc590005) is channels right-shifted by 12
bits -- i.e. (address_of_field) >> 12. That's not a coincidence: it
is the exact shape of glibc's "safe-linking" freelist obfuscation
(protected_ptr = (&field >> 12) XOR next; with next == NULL this
collapses to &field >> 12). In other words, the chunk backing this
channels array appears to have already been free()'d by the time
channelLoop()'s thread read it, and what's being read back is
leftover allocator bookkeeping, not a real ssh_channel*.
--------------------------------------------------------------------
CONFIRMED LIVE UNDER VALGRIND (--trace-children=yes)
--------------------------------------------------------------------
Running with:
valgrind --trace-children=yes --tool=memcheck --track-origins=yes \
x2goclient --debug --libssh-debug
caught the same class of fault live:
==372207== Invalid read of size 8
==372207== at 0x4B23B9A: ssh_event_add_session (in libssh.so.4.10.5)
==372207== by 0x4B0927E: ssh_select (in libssh.so.4.10.5)
==372207== by 0x1FB638: ??? (in /usr/libexec/x2goclient)
==372207== by 0x1FF303: ??? (in /usr/libexec/x2goclient)
==372207== by 0x5C3AD20: ??? (in libQt5Core.so.5.15.15)
==372207== by ...start_thread / clone
==372207== Address 0x2a000005c8 is not stack'd, malloc'd or (recently)
free'd
...
==372207== Process terminating with default action of signal 11
(SIGSEGV): dumping core
(offsets 0x1FB638/0x1FF303 line up with channelLoop()/run() from the
gdb trace above.) Note this crash happened in a forked child process
running the x2goclient binary image itself (a separate valgrind
==PID== group, not ssh/nxproxy), consistent with x2goclient forking
off a second SshMasterConnection worker for the jump-host hop.
Full (redacted, de-duplicated) valgrind log is attached separately:
x2go-valgrind.sanitized.log
--------------------------------------------------------------------
POSSIBLE CONTRIBUTING FACTOR: SCP-SESSION RETRY LOOP
--------------------------------------------------------------------
The valgrind run that caught the crash also logged 26,670 repeated
attempts of:
x2go-DEBUG-../src/sshmasterconnection.cpp:1850> Error initializing
SCP session: Failed to open channel for scp
in a tight loop, right before the crash. This looks like a separate,
unbounded retry loop (worth its own bug report?) -- but it's also
plausibly what's making the race easy to hit: rapid repeated channel
creation/teardown is exactly the kind of hammering that would narrow
the timing window needed to trigger a use-after-free like this.
--------------------------------------------------------------------
PRIOR RELATED REPORTS
--------------------------------------------------------------------
This isn't the first time SshMasterConnection's threading has
produced crashes in this area -- see e.g.:
#187 "Destroying SshMasterConnection off the main thread leads to
a diagnostic from Qt debug libs"
#1460 "Windows client crashes if Jumphost runs NetBSD 6 (probably
race)"
Both point at the same general area (channel/session lifecycle vs.
thread boundaries) without a definitive fix landing, as far as I can
tell from the bug history.
--------------------------------------------------------------------
ATTACHMENT
--------------------------------------------------------------------
x2go-valgrind.sanitized.log -- de-duplicated, redacted valgrind/
memcheck log (hostnames, usernames, an internal IP, and what looked
like a live session cookie have been replaced with placeholders)
--Philipp with Claude 5 Sonnet
[x2go-valgrind.sanitized.log (text/x-log, attachment)]
Information forwarded
to x2go-dev@lists.x2go.org, X2Go Developers <x2go-dev@lists.x2go.org>:
Bug#1637; Package x2goclient.
(Mon, 21 Sep 2026 09:20:01 GMT) (full text, mbox, link).
Acknowledgement sent
to Philipp Klenze <p.klenze@gsi.de>:
Extra info received and forwarded to list. Copy sent to X2Go Developers <x2go-dev@lists.x2go.org>.
(Mon, 21 Sep 2026 09:20:02 GMT) (full text, mbox, link).
Message #10 received at 1637@bugs.x2go.org (full text, mbox, reply):
[Message part 1 (text/plain, inline)]
Spend a few more tokens to fix this.
Note that I have not carefully verified the patch, but it seems to fix
both of the segfaults I encountered (but still does not make resume
work, that is a problem on the server side).
I also encountered an infinite loop (100% CPU usage while trying in vain
to start the proxy), but did not generate a patch for that yet, let me
know if you want one.
Cheers,
Philipp
--------------------------------------------------------------------
NOTE ON HOW THIS FOLLOW-UP WAS PRODUCED
--------------------------------------------------------------------
Same as the original report: debugging and patch drafting done
together with an LLM assistant (Claude) at my direction. I built,
ran, and verified everything myself. Flagging this again per the
original subject tag.
--------------------------------------------------------------------
SUMMARY
--------------------------------------------------------------------
Root cause found and fixed. The crash was NOT a libssh bug or a
generic race -- it was a x2goclient-side logic bug in
SshMasterConnection::createChannelConnection() / channelLoop(),
specifically a dangling-pointer bug in the array x2goclient hands to
ssh_select(). Two distinct issues in that area, both patched below.
I've been running the patched build since and no longer see this
crash, including in a case that previously crashed within seconds of
"resuming normal session".
--------------------------------------------------------------------
ROOT CAUSE #1: stale channel pointer not cleared on error
--------------------------------------------------------------------
In createChannelConnection() (src/sshmasterconnection.cpp), when
ssh_channel_open_session() or ssh_channel_request_exec() fails after
a channel has already been stored via
channelConnections[i].channel = channel;
the error paths do:
ssh_channel_free (channel);
...
return (false);
without resetting channelConnections[i].channel back to 0. The
caller in channelLoop() just does "continue;" on failure, so the
stale entry survives.
On the *next* pass, createChannelConnection() checks:
if ( channelConnections.at ( i ).channel==0l )
which is false (the pointer is non-null, just freed), so the
"create a new channel" branch is skipped entirely, and the dangling
pointer is handed straight to ssh_select() via read_chan[i]. Confirmed
with valgrind: an ssh_channel allocated and freed with both call
sites inside channelLoop() itself, then read again on the very next
ssh_select() pass.
Fix: explicitly reset channelConnections[i].channel = 0l on both
error paths, right after ssh_channel_free().
--------------------------------------------------------------------
ROOT CAUSE #2: uninitialized array slot on same-iteration failure
--------------------------------------------------------------------
Separately, in channelLoop():
ssh_channel* read_chan=new ssh_channel[channelConnections.size() +1];
ssh_channel* out_chan=new ssh_channel[channelConnections.size() +1];
read_chan[channelConnections.size() ]=NULL;
new T[n] does not zero-initialize a plain pointer array. For any
index whose createChannelConnection() call fails and returns early
(before reaching the "read_chan[i] = ..." line at the end of that
function), read_chan[i] is left as raw heap garbage, not NULL --
so ssh_select()'s NULL-terminated scan of the array walks straight
into it. This one can crash within a single channelLoop() iteration,
with no prior history needed, which is what I saw when a channel
failed immediately on session resume.
Fix: value-initialize both arrays ("new ssh_channel[n]()") so every
untouched slot starts NULL rather than garbage.
--------------------------------------------------------------------
PATCH
--------------------------------------------------------------------
Both fixes, against the 4.1.2.3 source tree (attached in full as
x2goclient-dangling-channel-v2.patch; apply with "patch -p0" from
the top of the source tree):
--- src/sshmasterconnection.cpp.orig
+++ src/sshmasterconnection.cpp
@@ -2205,6 +2205,11 @@
/* Free channel. */
ssh_channel_free (channel);
+ /* Local patch: without this, channelConnections[i].channel
+ * keeps pointing at freed memory, and the next loop
+ * iteration's "channel==0l" check wrongly treats it as
+ * still valid, handing a dangling pointer straight to
+ * ssh_select(). */
+ channelConnections[i].channel = 0l;
emit ioErr ( channelConnections[i].creator, errorMsg,
err );
x2goDebug<<errorMsg.left (errorMsg.size () - 1)<<":
"<<err<<endl;
@@ -2219,6 +2224,8 @@
/* Close connection and free channel. */
ssh_channel_close (channel);
ssh_channel_free (channel);
+ /* Local patch: see comment above. */
+ channelConnections[i].channel = 0l;
emit ioErr ( channelConnections[i].creator, errorMsg,
err );
x2goDebug<<errorMsg.left (errorMsg.size () - 1)<<":
"<<err<<endl;
@@ -1959,8 +1966,14 @@
}
- ssh_channel* read_chan=new
ssh_channel[channelConnections.size() +1];
- ssh_channel* out_chan=new ssh_channel[channelConnections.size()
+1];
+ /* Local patch: value-initialize ("()") so every slot starts NULL.
+ * createChannelConnection() can return early (on failure) before
+ * it reaches the line that fills in read_chan[i]/out_chan[i];
+ * without zero-init that slot is uninitialized garbage, not
+ * NULL, and ssh_select() walks off the end of it looking for a
+ * NULL terminator -- causing exactly the SIGSEGV in ssh_select()
+ * this works around. */
+ ssh_channel* read_chan=new
ssh_channel[channelConnections.size() +1]();
+ ssh_channel* out_chan=new ssh_channel[channelConnections.size()
+1]();
read_chan[channelConnections.size() ]=NULL;
(Line numbers are approximate/context-shifted between the two hunks
above; the attached patch file has exact, applicable hunks.)
--------------------------------------------------------------------
CAVEAT / OPEN QUESTION FOR MAINTAINERS
--------------------------------------------------------------------
My fix for issue #2 (zero-init) stops the crash but isn't fully
correct: ssh_select() stops at the first NULL it finds in the
channels[] array. If channel i fails and is left NULL while later
channels i+1, i+2, ... succeeded, this round's ssh_select() call
silently only covers channels 0..i-1 -- the rest are skipped for
that ~500ms iteration (self-corrects next iteration, since the array
is rebuilt fresh each pass, but it's not truly "correct"). A proper
fix would compact the array to skip failed slots rather than
truncate at them. I kept my patch minimal/local rather than
attempting that restructuring myself.
--------------------------------------------------------------------
NEXT BUG: INFINITE LOOP ON CONNECTION FAILURE
--------------------------------------------------------------------
Separately, while testing the fix above under a failure condition (the
server-side issue described below), I caught a related but distinct
problem in the same function: when ssh_channel_open_session() keeps
failing (e.g. the underlying SSH transport is still "connected" but the
server keeps rejecting the channel), channelLoop() has no failure cap or
backoff -- it just retries immediately, forever. With the crash fixed,
this now manifests as a thread spinning at 100% CPU indefinitely, doing
real cryptographic work each iteration (confirmed via gdb: the hot
thread was consistently inside ssh_channel_open_session(), not an empty
spin), with no user-visible error and no give-up condition. I haven't
attempted a patch for this one since it's more of a design/policy gap
(how many retries, what backoff, how to surface the eventual failure to
the user) than a straightforward bug fix, and I'd rather get a read on
whether the two patches above are wanted before writing more.
--------------------------------------------------------------------
UNRELATED FINDING WHILE DEBUGGING (separate report, x2goserver not
x2goclient)
--------------------------------------------------------------------
While chasing why sessions still failed to connect after this fix,
I found a real, reproducible bug on the server side, in x2goserver
(not x2goclient) -- filing this separately, but noting it here since
it's what made the client-side race easy to trigger in the first
place (rapid failed reconnect/channel-creation attempts):
/usr/lib/x2go/x2gormport line 34:
my $port=shift or die;
"shift or die" is falsy-checked, not definedness-checked, so this
(and two equivalent lines further down the call chain, in
X2Go::Server::DB::db_rmport and the SQLite3 backend's db_rmport)
dies whenever a legitimately-zero port value is passed -- which
happens routinely for ports not previously in use by a session.
Reproduced directly server-side:
$ x2goresume-session pklenze-58-... 1920x1080 adsl 16m-jpeg-9 us
pc105/us 1 both no
Died at /usr/lib/x2go/x2gormport line 34.
Died at /usr/lib/x2go/x2gormport line 34.
gr_port=34759
sound_port=34760
fs_port=34761
Connection to lxi111 closed by remote host.
The resume script prints the port numbers regardless of whether
x2gormport actually succeeded, so the client is told to forward to
a port whose server-side setup silently failed -- producing
"Channel opening failure ... Connection refused" client-side, which
is what started this whole thread. I'll file this against x2goserver
separately once I've confirmed the fix; flagging here for context /
in case it's useful to anyone hitting a similar "resume looks fine
but graphics never connects" symptom.
--------------------------------------------------------------------
ATTACHMENT
--------------------------------------------------------------------
x2goclient-dangling-channel-v2.patch -- full patch, apply with
"patch -p0" from the top of the x2goclient-4.1.2.3 source tree.
[x2goclient-dangling-channel-v2.patch (text/x-patch, attachment)]
Send a report that this bug log contains spam.
Debbugs is free software and licensed under the terms of the GNU Public License version 2. The current version can be obtained from https://bugs.debian.org/debbugs-source/.
Copyright © 1999 Darren O. Benham, 1997,2003 nCipher Corporation Ltd, 1994-97 Ian Jackson, 2005-2017 Don Armstrong, and many other contributors.