public inbox for io-uring@vger.kernel.org
 help / color / mirror / Atom feed
* [PATCH v3 00/17] coredump & signals: an impossible affair
@ 2026-09-21 13:44 Christian Brauner
  2026-09-21 13:44 ` [PATCH v3 01/17] coredump: hold RCU while releasing parked threads Christian Brauner
                   ` (16 more replies)
  0 siblings, 17 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:44 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable),
	stable

Hey,

I asked Chris to look at the coredump code with kres and it found a few
bugs. I started looking as well and found a few more. Here's a fixes
series. I also used TLA+ modeling for this.

Fixes in here:

- UAF in coredump_finish(): a parked thread can be freed before it is
  woken
- only SIGKILL and the freezers interrupt a dump now, cgroup v2 included
- core_pattern is parsed from a snapshot instead of racing the sysctl
- the io-wq exit bit wasn't ordered against worker creation task work,
  exit could hang on worker_done
- shared signals are no longer retargeted to the dumper, that truncated
  cores
- a failed fork released its files under scx_fork_rwsem, deadlock
- a session leader's exit lost the SIGHUP for the foreground job
- an io-wq worker of an SQPOLL ring as the dumper deadlocks the group,
  user workers never dump now
- PTRACE_SETSIGMASK can't unmask a user worker anymore, the only way in
- exec cancels io_uring before de_thread(), nothing adds a thread after
  it
- no io threads from PF_SIGNALED or PF_POSTCOREDUMP creators, they
  broke threads_remaining
- descriptor tables are closed highest fd first again, the order the
  deferred puts had

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
Changes in v3:
- Address Oleg's reviews.
- Add a couple more fixes.
- Link to v2: https://patch.msgid.link/20260917-work-coredump-fixes-v2-0-f3787fcda051@kernel.org

Changes in v2:
- Add fixes for retarget_shared_signal().
- Expand fixes for signal_pending().
- Link to v1: https://patch.msgid.link/20260915-work-coredump-fixes-v1-0-f354ca41780c@kernel.org

---
Christian Brauner (18):
      Merge patch series "files: make closing files synchronous for close_range(), exec, exit"
      coredump: hold RCU while releasing parked threads
      signal: only SIGKILL and the freezers interrupt a coredumping task
      coredump: parse a snapshot of core_pattern
      io-wq: order the exit bit against worker creation task work
      signal: don't retarget shared signals in a dying thread group
      selftests/coredump: test shared signal retargeting during a dump
      fork: release the files of a failed fork after sched_cancel_fork()
      exit: hang up the tty before closing the files
      ptrace: refuse to change the signal mask of a user worker
      selftests/coredump: test a user worker as the coredumping thread
      selftests/coredump: expect PTRACE_SETSIGMASK to be refused on a user worker
      exec: cancel io_uring requests before de_thread()
      fork: move the coredump and exec checks into create_io_thread()
      fork: don't create io threads once PF_POSTCOREDUMP is set
      fork: use SIG_KERNEL_ONLY_MASK for the user worker signal mask
      signal: enforce the user worker signal mask in __set_task_blocked()
      fs: close files from the highest descriptor down

 fs/coredump.c                                      |  73 ++--
 fs/exec.c                                          |  11 +-
 fs/file.c                                          |  55 +--
 include/linux/sched/signal.h                       |  19 +-
 io_uring/io-wq.c                                   |   2 +
 kernel/exit.c                                      |   5 +-
 kernel/fork.c                                      |  20 +-
 kernel/ptrace.c                                    |   6 +
 kernel/signal.c                                    |  22 +
 tools/testing/selftests/coredump/.gitignore        |   2 +
 tools/testing/selftests/coredump/Makefile          |   6 +-
 .../selftests/coredump/coredump_signal_test.c      | 238 +++++++++++
 .../selftests/coredump/coredump_worker_test.c      | 447 +++++++++++++++++++++
 13 files changed, 832 insertions(+), 74 deletions(-)
---
base-commit: dadceac9d20a1c90269aafd7746f7c0c05879ad3
change-id: 20260915-work-coredump-fixes-edf98c80fe78


^ permalink raw reply	[flat|nested] 29+ messages in thread

* [PATCH v3 01/17] coredump: hold RCU while releasing parked threads
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
@ 2026-09-21 13:44 ` Christian Brauner
  2026-09-21 13:44 ` [PATCH v3 02/17] signal: only SIGKILL and the freezers interrupt a coredumping task Christian Brauner
                   ` (15 subsequent siblings)
  16 siblings, 0 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:44 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable),
	stable

coredump_finish() releases the parked threads by clearing ->task in
their core_thread entry and calling wake_up_process() on each of them.
It does that without holding a reference on the task and without being
inside rcu.

But calling wake_up_process(task) without rcu here isn't safe. A parked
thread doesn't need that wakeup to leave. A spurious wakeup or a
preemption after the store is enough for the task to go away.

So if the coredump client is preempted between the store and
wake_up_process() the thread can exit and be freed in the meantime and
try_to_wake_up() takes pi_lock in freed memory.

Hold rcu across the loop.

Fixes: a94e2d408eae ("coredump: kill mm->core_done")
Cc: stable@vger.kernel.org
Acked-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/coredump.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/fs/coredump.c b/fs/coredump.c
index 16b331b686fb..6c0c597ec324 100644
--- a/fs/coredump.c
+++ b/fs/coredump.c
@@ -561,6 +561,8 @@ static void coredump_finish(enum coredump_state state)
 	current->signal->core_state = NULL;
 	spin_unlock_irq(&current->sighand->siglock);
 
+	/* A released thread may exit and be freed before it is woken. */
+	guard(rcu)();
 	while ((curr = next) != NULL) {
 		next = curr->next;
 		task = curr->task;
@@ -569,6 +571,7 @@ static void coredump_finish(enum coredump_state state)
 		 * ->task == NULL before we read ->next.
 		 */
 		smp_mb();
+		/* Any wakeup now lets the thread exit, rcu keeps it alive. */
 		curr->task = NULL;
 		wake_up_process(task);
 	}

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 02/17] signal: only SIGKILL and the freezers interrupt a coredumping task
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
  2026-09-21 13:44 ` [PATCH v3 01/17] coredump: hold RCU while releasing parked threads Christian Brauner
@ 2026-09-21 13:44 ` Christian Brauner
  2026-09-22 12:55   ` Oleg Nesterov
  2026-09-21 13:44 ` [PATCH v3 03/17] coredump: parse a snapshot of core_pattern Christian Brauner
                   ` (14 subsequent siblings)
  16 siblings, 1 reply; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:44 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable),
	stable

For quite a while now coredump has allowed either SIGKILL or freezing to
interrupt an ongoing coredump. Since we seem to like it complicated
"freezing" can mean a lot of things:

(1) Power Management induced freezing
(2) cgroup v1 freezer controller induced freezing
(3) cgroup v2 freezer controller induced freezing

So (1) and (2) are handled by checking freezing() but (3) isn't.
(3) uses cgroup_freeze_task() which sets JOBCTL_TRAP_FREEZE and calls
signal_wake_up() on every task in the cgroup. The trap isn't handled by
the coredump code.

Which means the current logic is inconsistent. Power and cgroup v1
freeze and abort the coredump. With cgroup v2 it depends on where the
core goes. A dump to a file completes and the cgroup v2 freeze has to
wait for it to finish. When dumping to a pipe or socket the dump is
interrupted at the next write because of trap's TIF_SIGPENDING. Which is
it? Let's be consistent and align (1)-(3): freezing interrupts an
ongoing coredump and doesn't make it wait until the dump is done.

So make signal_pending() report only SIGKILL, freezing() and the cgroup
v2 freezer trap for a task that has PF_DUMPCORE set. freezing() lives in
linux/freezer.h and is a static branch. So keep that check a separate
helper instead of in signal_pending() itself.

Make dump_interrupted() check for JOBCTL_TRAP_FREEZE as well so a cgroup
v2 freeze keeps aborting the dump like PM and cgroup v1 do.

Fixes: 403bad72b67d ("coredump: only SIGKILL should interrupt the coredumping task")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/coredump.c                | 11 ++++++-----
 include/linux/sched/signal.h | 19 +++++++++++++------
 kernel/signal.c              |  8 ++++++++
 3 files changed, 27 insertions(+), 11 deletions(-)

diff --git a/fs/coredump.c b/fs/coredump.c
index 6c0c597ec324..40eca2b85b81 100644
--- a/fs/coredump.c
+++ b/fs/coredump.c
@@ -580,12 +580,13 @@ static void coredump_finish(enum coredump_state state)
 static bool dump_interrupted(void)
 {
 	/*
-	 * SIGKILL or freezing() interrupt the coredumping. Perhaps we
-	 * can do try_to_freeze() and check __fatal_signal_pending(),
-	 * but then we need to teach dump_write() to restart and clear
-	 * TIF_SIGPENDING.
+	 * SIGKILL, freezing() or the cgroup v2 freezer trap interrupt the
+	 * coredumping. Perhaps we can do try_to_freeze() and check
+	 * __fatal_signal_pending(), but then we need to teach dump_write()
+	 * to restart and clear TIF_SIGPENDING.
 	 */
-	return fatal_signal_pending(current) || freezing(current);
+	return fatal_signal_pending(current) || freezing(current) ||
+	       (READ_ONCE(current->jobctl) & JOBCTL_TRAP_FREEZE);
 }
 
 static void wait_for_dump_helpers(struct file *file)
diff --git a/include/linux/sched/signal.h b/include/linux/sched/signal.h
index e039e29cd8c5..3835fbf7d80f 100644
--- a/include/linux/sched/signal.h
+++ b/include/linux/sched/signal.h
@@ -384,6 +384,13 @@ static inline int task_sigpending(struct task_struct *p)
 	return unlikely(test_tsk_thread_flag(p,TIF_SIGPENDING));
 }
 
+static inline int __fatal_signal_pending(struct task_struct *p)
+{
+	return unlikely(sigismember(&p->pending.signal, SIGKILL));
+}
+
+bool coredump_signal_pending(struct task_struct *p);
+
 static inline int signal_pending(struct task_struct *p)
 {
 	/*
@@ -393,12 +400,12 @@ static inline int signal_pending(struct task_struct *p)
 	 */
 	if (unlikely(test_tsk_thread_flag(p, TIF_NOTIFY_SIGNAL)))
 		return 1;
-	return task_sigpending(p);
-}
-
-static inline int __fatal_signal_pending(struct task_struct *p)
-{
-	return unlikely(sigismember(&p->pending.signal, SIGKILL));
+	if (!task_sigpending(p))
+		return 0;
+	/* A coredumping task only stops for SIGKILL or the freezer. */
+	if (unlikely(READ_ONCE(p->flags) & PF_DUMPCORE))
+		return coredump_signal_pending(p);
+	return 1;
 }
 
 static inline int fatal_signal_pending(struct task_struct *p)
diff --git a/kernel/signal.c b/kernel/signal.c
index ec30550951ec..d8bb1f168055 100644
--- a/kernel/signal.c
+++ b/kernel/signal.c
@@ -717,6 +717,14 @@ void signal_wake_up_state(struct task_struct *t, unsigned int state)
 		kick_process(t);
 }
 
+/* Only SIGKILL or a freezer interrupt a coredump, see dump_interrupted(). */
+bool coredump_signal_pending(struct task_struct *p)
+{
+	return __fatal_signal_pending(p) || freezing(p) ||
+	       (READ_ONCE(p->jobctl) & JOBCTL_TRAP_FREEZE);
+}
+EXPORT_SYMBOL(coredump_signal_pending);
+
 static inline void posixtimer_sig_ignore(struct task_struct *tsk, struct sigqueue *q);
 
 static void sigqueue_free_ignored(struct task_struct *tsk, struct sigqueue *q)

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 03/17] coredump: parse a snapshot of core_pattern
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
  2026-09-21 13:44 ` [PATCH v3 01/17] coredump: hold RCU while releasing parked threads Christian Brauner
  2026-09-21 13:44 ` [PATCH v3 02/17] signal: only SIGKILL and the freezers interrupt a coredumping task Christian Brauner
@ 2026-09-21 13:44 ` Christian Brauner
  2026-09-21 13:44 ` [PATCH v3 04/17] io-wq: order the exit bit against worker creation task work Christian Brauner
                   ` (13 subsequent siblings)
  16 siblings, 0 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:44 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable)

This is a long-standing problem that I discussed a while back with Jann.
I didn't care enough about it to really fix it and it's from the
before-fore-times.

coredump_parse() reads core_pattern directly while it can concurrently
be modified. So it reads the first byte, figures out what mode is
wanted, then allocates the number buffer and then consumes the rest of
the core_pattern string.

Say the sysctl handler updates the core_pattern array byte by byte
(idiotic but supported). So that can lead to all kinds of insane mixups.
Say you could transform the old "|/usr/bin/helper" and the new
"/tmp/core.%p" into a usermodehelper started as "tmp/core.<pid>".

So copy what proc_do_uts_string() does and let the handler run
proc_dostring() on a copy, validate the copy and make it visible beneath
a spinlock. Then coredump_parse() can take a snapshot under the same
spinlock and parse a stable copy.

From now on, rejected patterns are never visible and we can drop the
whole rollback logic. It has the same minor defect that utsname has,
namely that two writers can race on a non-zero offset. Irrelevant imho.

Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/coredump.c | 59 +++++++++++++++++++++++++++++++++++++----------------------
 1 file changed, 37 insertions(+), 22 deletions(-)

diff --git a/fs/coredump.c b/fs/coredump.c
index 40eca2b85b81..d5d76704df81 100644
--- a/fs/coredump.c
+++ b/fs/coredump.c
@@ -85,6 +85,8 @@ static int core_uses_pid;
 static unsigned int core_pipe_limit;
 static unsigned int core_sort_vma;
 static char core_pattern[CORENAME_MAX_SIZE] = "core";
+/* Taken around every copy in and out of core_pattern. */
+static DEFINE_SPINLOCK(core_pattern_lock);
 static int core_name_size = CORENAME_MAX_SIZE;
 unsigned int core_file_note_size_limit = CORE_FILE_NOTE_SIZE_DEFAULT;
 static atomic_t core_pipe_count = ATOMIC_INIT(0);
@@ -240,11 +242,16 @@ static bool coredump_parse(struct core_name *cn, struct coredump_params *cprm,
 			   size_t **argv, int *argc)
 {
 	const struct cred *cred = current_cred();
-	const char *pat_ptr = core_pattern;
+	char pattern[CORENAME_MAX_SIZE];
+	const char *pat_ptr = pattern;
 	bool was_space = false;
 	int pid_in_pattern = 0;
 	int err = 0;
 
+	/* The sysctl handler may be publishing a new pattern. */
+	scoped_guard(spinlock, &core_pattern_lock)
+		strscpy(pattern, core_pattern);
+
 	cprm->mask = COREDUMP_KERNEL;
 	if (core_pipe_limit)
 		cprm->mask |= COREDUMP_WAIT;
@@ -1640,11 +1647,11 @@ void validate_coredump_safety(void)
 	}
 }
 
-static inline bool check_coredump_socket(void)
+static inline bool check_coredump_socket(const char *pattern)
 {
 	const char *p;
 
-	if (core_pattern[0] != '@')
+	if (pattern[0] != '@')
 		return true;
 
 	/*
@@ -1656,16 +1663,16 @@ static inline bool check_coredump_socket(void)
 		return false;
 
 	/* Must be an absolute path... */
-	if (core_pattern[1] != '/') {
+	if (pattern[1] != '/') {
 		/* ... or the socket request protocol... */
-		if (core_pattern[1] != '@')
+		if (pattern[1] != '@')
 			return false;
 		/* ... and if so must be an absolute path. */
-		if (core_pattern[2] != '/')
+		if (pattern[2] != '/')
 			return false;
-		p = &core_pattern[2];
+		p = &pattern[2];
 	} else {
-		p = &core_pattern[1];
+		p = &pattern[1];
 	}
 
 	/* The path obviously cannot exceed UNIX_PATH_MAX. */
@@ -1673,7 +1680,7 @@ static inline bool check_coredump_socket(void)
 		return false;
 
 	/* Must not contain ".." in the path. */
-	if (name_contains_dotdot(core_pattern))
+	if (name_contains_dotdot(pattern))
 		return false;
 
 	return true;
@@ -1682,27 +1689,35 @@ static inline bool check_coredump_socket(void)
 static int proc_dostring_coredump(const struct ctl_table *table, int write,
 		  void *buffer, size_t *lenp, loff_t *ppos)
 {
+	char pattern[CORENAME_MAX_SIZE];
+	const struct ctl_table tmp = {
+		.procname	= table->procname,
+		.data		= pattern,
+		.maxlen		= sizeof(pattern),
+	};
+	bool changed = false;
 	int error;
-	ssize_t retval;
-	char old_core_pattern[CORENAME_MAX_SIZE];
-
-	if (!write)
-		return proc_dostring(table, write, buffer, lenp, ppos);
 
-	retval = strscpy(old_core_pattern, core_pattern, CORENAME_MAX_SIZE);
+	/* Work on a copy, proc_dostring() appends at *ppos. */
+	scoped_guard(spinlock, &core_pattern_lock)
+		strscpy(pattern, core_pattern);
 
-	error = proc_dostring(table, write, buffer, lenp, ppos);
-	if (error)
+	error = proc_dostring(&tmp, write, buffer, lenp, ppos);
+	if (error || !write)
 		return error;
 
-	if (!check_coredump_socket()) {
-		strscpy(core_pattern, old_core_pattern, retval + 1);
+	if (!check_coredump_socket(pattern))
 		return -EINVAL;
-	}
 
-	if (strncmp(old_core_pattern, core_pattern, CORENAME_MAX_SIZE))
+	/* Publish the validated pattern whole. */
+	scoped_guard(spinlock, &core_pattern_lock) {
+		changed = strncmp(pattern, core_pattern, CORENAME_MAX_SIZE);
+		if (changed)
+			strscpy(core_pattern, pattern);
+	}
+	if (changed)
 		validate_coredump_safety();
-	return error;
+	return 0;
 }
 
 static const unsigned int core_file_note_size_min = CORE_FILE_NOTE_SIZE_DEFAULT;

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 04/17] io-wq: order the exit bit against worker creation task work
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (2 preceding siblings ...)
  2026-09-21 13:44 ` [PATCH v3 03/17] coredump: parse a snapshot of core_pattern Christian Brauner
@ 2026-09-21 13:44 ` Christian Brauner
  2026-09-21 13:44 ` [PATCH v3 05/17] signal: don't retarget shared signals in a dying thread group Christian Brauner
                   ` (12 subsequent siblings)
  16 siblings, 0 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:44 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable),
	stable

io_wq_exit_workers() cancels task work for queued worker creation with
task_work_cancel_match(). That has a plain load of task->task_works.

io_queue_worker_create() adds the new entry and tests IO_WQ_BIT_EXIT to
cancel it if the workqueue is already on its way out.

That must be ordered. task_work_add() has a full barrier via cmpxchg().
But io_wq_exit_start() sets the bit with set_bit() which doesn't have
any memory ordering.

So afaict, on weakly ordered architectures the exiting task may load
task_works before the store of the bit is visible. The other side tests
the bit before the store is visible as well. So the entry remains queued
with worker_refs and the exiting task waits on worker_done indefinitely.

If the task still has rings then io_uring_del_tctx_node() provides the
barrier via test_and_set_bit() in io_wq_set_exit_on_idle(). When it
has closed all rings though that barrier is gone.

Add the barrier after the set_bit().

Fixes: 71a85387546e ("io-wq: check for wq exit after adding new worker task_work")
Cc: stable@vger.kernel.org
Reviewed-by: Jens Axboe <axboe@kernel.dk>
Acked-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 io_uring/io-wq.c | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/io_uring/io-wq.c b/io_uring/io-wq.c
index 2ca223e47d41..2a980e86dd94 100644
--- a/io_uring/io-wq.c
+++ b/io_uring/io-wq.c
@@ -1324,6 +1324,8 @@ static bool io_task_work_match(struct callback_head *cb, void *data)
 void io_wq_exit_start(struct io_wq *wq)
 {
 	set_bit(IO_WQ_BIT_EXIT, &wq->state);
+	/* Pairs with task_work_add() in io_queue_worker_create(). */
+	smp_mb__after_atomic();
 }
 
 static void io_wq_cancel_tw_create(struct io_wq *wq)

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 05/17] signal: don't retarget shared signals in a dying thread group
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (3 preceding siblings ...)
  2026-09-21 13:44 ` [PATCH v3 04/17] io-wq: order the exit bit against worker creation task work Christian Brauner
@ 2026-09-21 13:44 ` Christian Brauner
  2026-09-21 13:44 ` [PATCH v3 06/17] selftests/coredump: test shared signal retargeting during a dump Christian Brauner
                   ` (11 subsequent siblings)
  16 siblings, 0 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:44 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable),
	stable

If the calling thread has blocked a signal then
retarget_shared_pending() slingshots a shared pending signal to a
sibling thread. Today it only skips PF_EXITING threads but not a
coredumping thread.

But if SIGNAL_GROUP_EXIT is set no thread in the thread-group can
dequeue such signals anymore as get_signal() sends SIGKILL to every
thread and prepare_signal() drops any other signal.

exit_signals() also returns early on SIGNAL_GROUP_EXIT. But neither
sigprocmask() nor the mask restore in sigtimedwait() do.

Say a thread sleeps in sigtimedwait() and a sibling thread crashes. The
sleeping thread gets woken by the zap call. It restores its mask on the
way out and retargets a signal that was queued for the coredumping
process. That means TIF_SIGPENDING is set on the coredumping task which
zap_threads() had cleared. So a blocking write sees TIF_SIGPENDING and
truncates the dump.

Stop retargeting signals when SIGNAL_GROUP_EXIT is set and avoid
needlessly truncating coredumps.

Fixes: e6fa16ab9c1e ("signal: sigprocmask() should do retarget_shared_pending()")
Cc: stable@vger.kernel.org
Acked-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 kernel/signal.c | 4 ++++
 1 file changed, 4 insertions(+)

diff --git a/kernel/signal.c b/kernel/signal.c
index d8bb1f168055..c3bad983dd37 100644
--- a/kernel/signal.c
+++ b/kernel/signal.c
@@ -3117,6 +3117,10 @@ static void retarget_shared_pending(struct task_struct *tsk, sigset_t *which)
 	sigset_t retarget;
 	struct task_struct *t;
 
+	/* Nobody dequeues them in a dying group, see get_signal(). */
+	if (tsk->signal->flags & SIGNAL_GROUP_EXIT)
+		return;
+
 	sigandsets(&retarget, &tsk->signal->shared_pending.signal, which);
 	if (sigisemptyset(&retarget))
 		return;

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 06/17] selftests/coredump: test shared signal retargeting during a dump
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (4 preceding siblings ...)
  2026-09-21 13:44 ` [PATCH v3 05/17] signal: don't retarget shared signals in a dying thread group Christian Brauner
@ 2026-09-21 13:44 ` Christian Brauner
  2026-09-21 13:44 ` [PATCH v3 07/17] fork: release the files of a failed fork after sched_cancel_fork() Christian Brauner
                   ` (10 subsequent siblings)
  16 siblings, 0 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:44 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable)

A shared signal that is still queued when a thread crashes gets
retargeted to the dumper when a zapped sibling restores its signal mask
on the way out of sigtimedwait(). That sets TIF_SIGPENDING on the dumper
and cuts the core short at the next full pipe or socket buffer. Cover
that:

- one sibling sleeps in sigtimedwait() for SIGUSR1

- a vfork() child queues SIGUSR1 for the main thread while it sleeps
  killably and then the SIGSEGV that is dequeued first

- the coredump server holds its read back so the dumper blocks on the
  full socket buffer

Check that the core covers every segment with and without the queued
signal.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 tools/testing/selftests/coredump/.gitignore        |   1 +
 tools/testing/selftests/coredump/Makefile          |   4 +-
 .../selftests/coredump/coredump_signal_test.c      | 238 +++++++++++++++++++++
 3 files changed, 242 insertions(+), 1 deletion(-)

diff --git a/tools/testing/selftests/coredump/.gitignore b/tools/testing/selftests/coredump/.gitignore
index 097f52db0be9..c198f2ca5872 100644
--- a/tools/testing/selftests/coredump/.gitignore
+++ b/tools/testing/selftests/coredump/.gitignore
@@ -2,3 +2,4 @@
 stackdump_test
 coredump_socket_test
 coredump_socket_protocol_test
+coredump_signal_test
diff --git a/tools/testing/selftests/coredump/Makefile b/tools/testing/selftests/coredump/Makefile
index dece1a31d561..d6cd9a7e9ac0 100644
--- a/tools/testing/selftests/coredump/Makefile
+++ b/tools/testing/selftests/coredump/Makefile
@@ -3,7 +3,8 @@ CFLAGS += -Wall -O0 -g $(KHDR_INCLUDES) $(TOOLS_INCLUDES)
 
 TEST_GEN_PROGS := stackdump_test \
 		  coredump_socket_test \
-		  coredump_socket_protocol_test
+		  coredump_socket_protocol_test \
+		  coredump_signal_test
 TEST_FILES := stackdump
 
 include ../lib.mk
@@ -11,3 +12,4 @@ include ../lib.mk
 $(OUTPUT)/stackdump_test: coredump_test_helpers.c
 $(OUTPUT)/coredump_socket_test: coredump_test_helpers.c
 $(OUTPUT)/coredump_socket_protocol_test: coredump_test_helpers.c
+$(OUTPUT)/coredump_signal_test: coredump_test_helpers.c
diff --git a/tools/testing/selftests/coredump/coredump_signal_test.c b/tools/testing/selftests/coredump/coredump_signal_test.c
new file mode 100644
index 000000000000..fdf48b144781
--- /dev/null
+++ b/tools/testing/selftests/coredump/coredump_signal_test.c
@@ -0,0 +1,238 @@
+// SPDX-License-Identifier: GPL-2.0
+
+#include <fcntl.h>
+#include <pthread.h>
+#include <signal.h>
+#include <sys/mman.h>
+#include <sys/socket.h>
+#include <sys/stat.h>
+#include <sys/syscall.h>
+#include <sys/un.h>
+#include <unistd.h>
+
+#include "coredump_test.h"
+
+/* Big enough to fill the socket buffer many times over. */
+#define CRASH_MAPPING_SIZE (32 * 1024 * 1024)
+
+FIXTURE_SETUP(coredump)
+{
+	FILE *file;
+	int ret;
+
+	self->pid_coredump_server = -ESRCH;
+	self->fd_tmpfs_detached = -1;
+	file = fopen("/proc/sys/kernel/core_pattern", "r");
+	ASSERT_NE(NULL, file);
+
+	ret = fread(self->original_core_pattern, 1, sizeof(self->original_core_pattern), file);
+	ASSERT_TRUE(ret || feof(file));
+	ASSERT_LT(ret, sizeof(self->original_core_pattern));
+
+	self->original_core_pattern[ret] = '\0';
+
+	ret = fclose(file);
+	ASSERT_EQ(0, ret);
+}
+
+FIXTURE_TEARDOWN(coredump)
+{
+	const char *reason;
+	FILE *file;
+	int ret, status;
+
+	if (self->pid_coredump_server > 0) {
+		kill(self->pid_coredump_server, SIGTERM);
+		waitpid(self->pid_coredump_server, &status, 0);
+	}
+	unlink("/tmp/coredump.file");
+	unlink("/tmp/coredump.socket");
+
+	file = fopen("/proc/sys/kernel/core_pattern", "w");
+	if (!file) {
+		reason = "Unable to open core_pattern";
+		goto fail;
+	}
+
+	ret = fprintf(file, "%s", self->original_core_pattern);
+	if (ret < 0) {
+		reason = "Unable to write to core_pattern";
+		goto fail;
+	}
+
+	ret = fclose(file);
+	if (ret) {
+		reason = "Unable to close core_pattern";
+		goto fail;
+	}
+
+	return;
+fail:
+	/* This should never happen */
+	fprintf(stderr, "Failed to cleanup coredump test: %s\n", reason);
+}
+
+static volatile int waiter_ready;
+
+static void usr1_handler(int sig)
+{
+}
+
+/* Sleeps in sigtimedwait() with SIGUSR1 unblocked only inside the kernel. */
+static void *sigwaiter(void *arg)
+{
+	sigset_t set;
+	siginfo_t info;
+
+	sigemptyset(&set);
+	sigaddset(&set, SIGUSR1);
+	__atomic_store_n(&waiter_ready, 1, __ATOMIC_RELEASE);
+	for (;;)
+		sigtimedwait(&set, &info, NULL);
+	return NULL;
+}
+
+/*
+ * Crash with a shared SIGUSR1 still queued for this thread. The vfork()
+ * child queues it while we sleep killably and then the SIGSEGV that is
+ * dequeued first. The zap wakes the sibling out of sigtimedwait() and its
+ * mask restore retargets SIGUSR1 to the dumper.
+ */
+static void crashing_child_retarget(bool queue_shared)
+{
+	sigset_t all, old;
+	pthread_t thread;
+	pid_t pid, tid;
+	char *p;
+
+	p = mmap(NULL, CRASH_MAPPING_SIZE, PROT_READ | PROT_WRITE,
+		 MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+	if (p == MAP_FAILED)
+		_exit(EXIT_FAILURE);
+	memset(p, 0x5a, CRASH_MAPPING_SIZE);
+
+	signal(SIGUSR1, usr1_handler);
+	sigfillset(&all);
+	pthread_sigmask(SIG_BLOCK, &all, &old);
+	/* One waiter only, retarget stops at the first thread not blocking it. */
+	if (pthread_create(&thread, NULL, sigwaiter, NULL))
+		_exit(EXIT_FAILURE);
+	pthread_sigmask(SIG_SETMASK, &old, NULL);
+	while (!__atomic_load_n(&waiter_ready, __ATOMIC_ACQUIRE))
+		usleep(1000);
+	usleep(50 * 1000);
+
+	pid = getpid();
+	tid = syscall(SYS_gettid);
+	if (vfork() == 0) {
+		if (queue_shared)
+			syscall(SYS_kill, pid, SIGUSR1);
+		syscall(SYS_tgkill, pid, tid, SIGSEGV);
+		syscall(SYS_exit, 0);
+	}
+
+	/* Not reached, the pending SIGSEGV dumps core. */
+	for (;;)
+		pause();
+}
+
+/*
+ * Dump the crashing child into /tmp/coredump.file through a server that
+ * holds the read back so the dumper blocks on the full socket buffer.
+ */
+static void run_coredump(struct __test_metadata *const _metadata,
+			 FIXTURE_DATA(coredump) *self, bool queue_shared)
+{
+	pid_t pid, pid_coredump_server;
+	int ipc_sockets[2];
+	int status;
+	char c;
+
+	unlink("/tmp/coredump.file");
+	unlink("/tmp/coredump.socket");
+	ASSERT_TRUE(set_core_pattern("@/tmp/coredump.socket"));
+	ASSERT_EQ(socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets), 0);
+
+	pid_coredump_server = fork();
+	ASSERT_GE(pid_coredump_server, 0);
+	if (pid_coredump_server == 0) {
+		int fd_server = -1, fd_coredump = -1, fd_core_file = -1;
+		int exit_code = EXIT_FAILURE;
+
+		close(ipc_sockets[0]);
+
+		fd_server = create_and_listen_unix_socket("/tmp/coredump.socket");
+		if (fd_server < 0)
+			goto out;
+
+		if (write_nointr(ipc_sockets[1], "1", 1) < 0)
+			goto out;
+		close(ipc_sockets[1]);
+
+		fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC);
+		if (fd_coredump < 0) {
+			fprintf(stderr, "%s: accept4 failed: %m\n", __func__);
+			goto out;
+		}
+
+		/* Let the dumper run into the full socket buffer first. */
+		sleep(1);
+
+		fd_core_file = creat("/tmp/coredump.file", 0644);
+		if (fd_core_file < 0) {
+			fprintf(stderr, "%s: creat failed: %m\n", __func__);
+			goto out;
+		}
+
+		if (recv_coredump_bytes(fd_coredump, fd_core_file) < 0)
+			goto out;
+
+		exit_code = EXIT_SUCCESS;
+out:
+		if (fd_core_file >= 0)
+			close(fd_core_file);
+		if (fd_coredump >= 0)
+			close(fd_coredump);
+		if (fd_server >= 0)
+			close(fd_server);
+		_exit(exit_code);
+	}
+	self->pid_coredump_server = pid_coredump_server;
+
+	EXPECT_EQ(close(ipc_sockets[1]), 0);
+	ASSERT_EQ(read_nointr(ipc_sockets[0], &c, 1), 1);
+	EXPECT_EQ(close(ipc_sockets[0]), 0);
+
+	pid = fork();
+	ASSERT_GE(pid, 0);
+	if (pid == 0)
+		crashing_child_retarget(queue_shared);
+
+	waitpid(pid, &status, 0);
+	ASSERT_TRUE(WIFSIGNALED(status));
+	ASSERT_EQ(WTERMSIG(status), SIGSEGV);
+	ASSERT_TRUE(WCOREDUMP(status));
+
+	wait_and_check_coredump_server(pid_coredump_server, _metadata, self);
+}
+
+static void check_coredump_complete(struct __test_metadata *const _metadata)
+{
+	int fd;
+
+	fd = open("/tmp/coredump.file", O_RDONLY | O_CLOEXEC);
+	ASSERT_GE(fd, 0);
+	ASSERT_TRUE(check_coredump_extent(fd));
+	close(fd);
+}
+
+TEST_F(coredump, retarget_shared_pending)
+{
+	run_coredump(_metadata, self, false);
+	check_coredump_complete(_metadata);
+
+	run_coredump(_metadata, self, true);
+	check_coredump_complete(_metadata);
+}
+
+TEST_HARNESS_MAIN

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 07/17] fork: release the files of a failed fork after sched_cancel_fork()
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (5 preceding siblings ...)
  2026-09-21 13:44 ` [PATCH v3 06/17] selftests/coredump: test shared signal retargeting during a dump Christian Brauner
@ 2026-09-21 13:44 ` Christian Brauner
  2026-09-21 13:44 ` [PATCH v3 08/17] exit: hang up the tty before closing the files Christian Brauner
                   ` (9 subsequent siblings)
  16 siblings, 0 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:44 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable)

Reorder copy_process() cleanup so sched_fork() taking scx_fork_rwsem
scx_pre_fork() is safe and isn't held around exiting files. Right now, a
fork that fails while another thread closed the bpf link fd will
deadlock against its own read side. It also blocks every fork on the
system:

    copy_process()                  holds scx_fork_rwsem for read
      exit_files()
        close_files()
          filp_close_sync()
            bpf_scx_unreg()
              kthread_flush_work()  waits for the disable work
                scx_root_disable()
                  percpu_down_write(&scx_fork_rwsem)

Simply release the child's files after sched_cancel_fork() dropped the
lock. The task starts out as a copy of its parent so p->files points to
the parent's table until copy_files() replaces it. exit_files() must not
run for a fork that failed before that. Clear p->files up front so
exit_files() is a no-op for those and make copy_files() set it
explicitly for CLONE_FILES.

Reported-by: Chris Mason <mason@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 kernel/fork.c | 9 ++++++---
 1 file changed, 6 insertions(+), 3 deletions(-)

diff --git a/kernel/fork.c b/kernel/fork.c
index 10be4a0ecb3f..50f5b3e2ca87 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -1676,6 +1676,7 @@ static int copy_files(u64 clone_flags, struct task_struct *tsk,
 
 	if (clone_flags & CLONE_FILES) {
 		atomic_inc(&oldf->count);
+		tsk->files = oldf;
 		return 0;
 	}
 
@@ -2199,6 +2200,8 @@ __latent_entropy struct task_struct *copy_process(
 	INIT_LIST_HEAD(&p->sibling);
 	rcu_copy_process(p);
 	p->vfork_done = NULL;
+	/* Set by copy_files(), exit_files() on the error path skips NULL. */
+	p->files = NULL;
 	spin_lock_init(&p->alloc_lock);
 
 	init_sigpending(&p->pending);
@@ -2300,7 +2303,7 @@ __latent_entropy struct task_struct *copy_process(
 		goto bad_fork_cleanup_semundo;
 	retval = copy_fs(clone_flags, p, args->umh);
 	if (retval)
-		goto bad_fork_cleanup_files;
+		goto bad_fork_cleanup_semundo;
 	retval = copy_sighand(clone_flags, p);
 	if (retval)
 		goto bad_fork_cleanup_fs;
@@ -2613,8 +2616,6 @@ __latent_entropy struct task_struct *copy_process(
 	__cleanup_sighand(p->sighand);
 bad_fork_cleanup_fs:
 	exit_fs(p); /* blocking */
-bad_fork_cleanup_files:
-	exit_files(p); /* blocking */
 bad_fork_cleanup_semundo:
 	exit_sem(p);
 bad_fork_cleanup_security:
@@ -2625,6 +2626,8 @@ __latent_entropy struct task_struct *copy_process(
 	perf_event_free_task(p);
 bad_fork_sched_cancel_fork:
 	sched_cancel_fork(p);
+	/* ->release() of a file may need scx_fork_rwsem for write. */
+	exit_files(p); /* blocking */
 bad_fork_cleanup_policy:
 	lockdep_free_task(p);
 #ifdef CONFIG_NUMA

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 08/17] exit: hang up the tty before closing the files
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (6 preceding siblings ...)
  2026-09-21 13:44 ` [PATCH v3 07/17] fork: release the files of a failed fork after sched_cancel_fork() Christian Brauner
@ 2026-09-21 13:44 ` Christian Brauner
  2026-09-24 12:10   ` Oleg Nesterov
  2026-09-21 13:44 ` [PATCH v3 09/17] ptrace: refuse to change the signal mask of a user worker Christian Brauner
                   ` (8 subsequent siblings)
  16 siblings, 1 reply; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:44 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable)

do_exit() closes the task's files in exit_files() and hangs up the
controlling tty of a session leader in disassociate_ctty(1) after
that. Since commit d99d38540bf0 ("fs: make close_files() synchronous")
the final __fput() of every file runs inside exit_files(). So when the
session leader holds the last open of its tty the tty is released
before disassociate_ctty() runs. tty_release() clears signal->tty for
the whole session in session_clear_tty() once the count drops to zero
and sends no signal doing so. disassociate_ctty(1) then finds neither a
tty nor a tty_old_pgrp and does nothing. The foreground process group
loses its SIGHUP:

    do_exit()
      exit_files()
        close_files()
          tty_release()             tty->count == 0
            session_clear_tty()     signal->tty = NULL, no signal
      disassociate_ctty(1)
        get_current_tty()           NULL
        signal->tty_old_pgrp        NULL, nothing sent

That only affects real ttys. For a pty the master's open keeps the
slave's count above zero. And it only affects a foreground job that
holds no descriptor to the tty anymore while its session leader exits.
Everything else is unchanged. The DTR drop on the last close happens in
tty_port_shutdown() regardless, stopped jobs get their SIGHUP from
kill_orphaned_pgrp() and signal->tty is cleared either way.

Before that commit the final __fput() ran from exit_task_work() which
comes after disassociate_ctty(). That order isn't old. Until v3.14
exit_task_work() came right after exit_files() and before v3.6 fput()
was synchronous, so the tty was always released first. Commit
c39df5fa37b0 ("exit: call disassociate_ctty() before
exit_task_namespaces()") moved disassociate_ctty() up to fix a pppd
crash and in front of exit_task_work() as a side effect. The hangup in
this case has worked since then and that's eleven years of userspace
being able to rely on it.

Hang the tty up before closing the files. This is the ordinary hangup
with the file still open: __tty_hangup() swaps in hung_up_tty_fops and
tty_release() runs from the close afterwards as it does when a modem
drops the line. disassociate_ctty() stays in front of
exit_task_namespaces() which the pppd fix needs.

Reported-by: Chris Mason <mason@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 kernel/exit.c | 5 +++--
 1 file changed, 3 insertions(+), 2 deletions(-)

diff --git a/kernel/exit.c b/kernel/exit.c
index 55dbea3b242e..9ed5eb03d0e1 100644
--- a/kernel/exit.c
+++ b/kernel/exit.c
@@ -1003,10 +1003,11 @@ void __noreturn do_exit(long code)
 
 	exit_sem(tsk);
 	exit_shm(tsk);
-	exit_files(tsk);
-	exit_fs(tsk);
+	/* Hang the tty up before the last close of it can clear the session. */
 	if (group_dead)
 		disassociate_ctty(1);
+	exit_files(tsk);
+	exit_fs(tsk);
 	exit_nsproxy_namespaces(tsk);
 	exit_task_work(tsk);
 	exit_thread(tsk);

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 09/17] ptrace: refuse to change the signal mask of a user worker
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (7 preceding siblings ...)
  2026-09-21 13:44 ` [PATCH v3 08/17] exit: hang up the tty before closing the files Christian Brauner
@ 2026-09-21 13:44 ` Christian Brauner
  2026-09-21 14:14   ` Oleg Nesterov
  2026-09-21 13:44 ` [PATCH v3 10/17] selftests/coredump: test a user worker as the coredumping thread Christian Brauner
                   ` (7 subsequent siblings)
  16 siblings, 1 reply; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:44 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable),
	stable

A user worker is created with every signal other than SIGKILL and
SIGSTOP blocked. It never runs user code and that mask is what keeps
get_signal() from dequeuing anything else for it. complete_signal()
never picks a thread that blocks the signal, tkill() queues on the
worker but nothing dequeues it. ptrace_signal() requeues an injected
signal that the tracee blocks.

PTRACE_SETSIGMASK is the only way to change that mask from the outside.
A tracer that clears it makes the worker eligible for every signal.
Whatever the worker then dequeues it can only act on by leaving.

Refuse PTRACE_SETSIGMASK for a user worker with -EPERM, the way
ptrace_attach() refuses a kernel thread. A worker is guaranteed to only
ever dequeue SIGKILL or SIGSTOP and an injected signal stays pending on
it, which is what already happens when the mask isn't touched.

Taking the request and quietly leaving the mask alone was the other
option, the way set_current_blocked() keeps SIGKILL and SIGSTOP
unblocked whatever userspace asks for. But then a tracer gets success
back with nothing changed and no way to tell that apart from a mask that
took effect.

Fixes: e8b33b8cfafc ("Revert "kernel: treat PF_IO_WORKER like PF_KTHREAD for ptrace/signals"")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 kernel/ptrace.c | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/kernel/ptrace.c b/kernel/ptrace.c
index d041645d9d17..4e9822a87aab 100644
--- a/kernel/ptrace.c
+++ b/kernel/ptrace.c
@@ -1227,6 +1227,12 @@ int ptrace_request(struct task_struct *child, long request,
 	case PTRACE_SETSIGMASK: {
 		sigset_t new_set;
 
+		/* A user worker only ever takes SIGKILL and SIGSTOP. */
+		if (child->flags & PF_USER_WORKER) {
+			ret = -EPERM;
+			break;
+		}
+
 		if (addr != sizeof(sigset_t)) {
 			ret = -EINVAL;
 			break;

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 10/17] selftests/coredump: test a user worker as the coredumping thread
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (8 preceding siblings ...)
  2026-09-21 13:44 ` [PATCH v3 09/17] ptrace: refuse to change the signal mask of a user worker Christian Brauner
@ 2026-09-21 13:44 ` Christian Brauner
  2026-09-21 13:45 ` [PATCH v3 11/17] selftests/coredump: expect PTRACE_SETSIGMASK to be refused on a user worker Christian Brauner
                   ` (6 subsequent siblings)
  16 siblings, 0 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:44 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable)

A tracer can clear the signal mask of an io-wq worker or an SQPOLL
thread with PTRACE_SETSIGMASK and inject a coredump signal. The worker
then runs vfs_coredump() itself. Cover that:

- a child keeps a ring, an idle io-wq worker and with SQPOLL the SQPOLL
  thread alive

- seize the chosen thread, stop it with SIGSTOP, clear its mask and
  detach with SIGSEGV

- require the thread group to be gone in bounded time, by the dump or
  by SIGKILL

The io-wq worker of an SQPOLL ring fails: the SQPOLL thread waits for
its workers to exit before it parks, the dumping worker waits for the
SQPOLL thread to park and the group is stuck in D state.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 tools/testing/selftests/coredump/.gitignore        |   1 +
 tools/testing/selftests/coredump/Makefile          |   4 +-
 .../selftests/coredump/coredump_worker_test.c      | 427 +++++++++++++++++++++
 3 files changed, 431 insertions(+), 1 deletion(-)

diff --git a/tools/testing/selftests/coredump/.gitignore b/tools/testing/selftests/coredump/.gitignore
index c198f2ca5872..e32f6e9006f6 100644
--- a/tools/testing/selftests/coredump/.gitignore
+++ b/tools/testing/selftests/coredump/.gitignore
@@ -3,3 +3,4 @@ stackdump_test
 coredump_socket_test
 coredump_socket_protocol_test
 coredump_signal_test
+coredump_worker_test
diff --git a/tools/testing/selftests/coredump/Makefile b/tools/testing/selftests/coredump/Makefile
index d6cd9a7e9ac0..57ec331f03a7 100644
--- a/tools/testing/selftests/coredump/Makefile
+++ b/tools/testing/selftests/coredump/Makefile
@@ -4,7 +4,8 @@ CFLAGS += -Wall -O0 -g $(KHDR_INCLUDES) $(TOOLS_INCLUDES)
 TEST_GEN_PROGS := stackdump_test \
 		  coredump_socket_test \
 		  coredump_socket_protocol_test \
-		  coredump_signal_test
+		  coredump_signal_test \
+		  coredump_worker_test
 TEST_FILES := stackdump
 
 include ../lib.mk
@@ -13,3 +14,4 @@ $(OUTPUT)/stackdump_test: coredump_test_helpers.c
 $(OUTPUT)/coredump_socket_test: coredump_test_helpers.c
 $(OUTPUT)/coredump_socket_protocol_test: coredump_test_helpers.c
 $(OUTPUT)/coredump_signal_test: coredump_test_helpers.c
+$(OUTPUT)/coredump_worker_test: coredump_test_helpers.c
diff --git a/tools/testing/selftests/coredump/coredump_worker_test.c b/tools/testing/selftests/coredump/coredump_worker_test.c
new file mode 100644
index 000000000000..81cadde0e752
--- /dev/null
+++ b/tools/testing/selftests/coredump/coredump_worker_test.c
@@ -0,0 +1,427 @@
+// SPDX-License-Identifier: GPL-2.0
+
+/*
+ * A user worker as the coredumping thread.
+ *
+ * io-wq workers and SQPOLL threads are threads of the process that never
+ * return to userspace. They block every signal but SIGKILL and SIGSTOP,
+ * but a tracer can replace that mask with PTRACE_SETSIGMASK and inject a
+ * coredump signal. get_signal() then runs vfs_coredump() in the worker.
+ * The worker's own exit bookkeeping runs only after the dump, so a
+ * zapped sibling that waits for it in its exit path deadlocks with the
+ * dumper and the whole thread group is stuck in D state.
+ *
+ * Inject SIGSEGV into a chosen thread and require that the thread group
+ * is gone in bounded time, either because the dump completed or because
+ * SIGKILL still works. A failure leaves the stuck process behind.
+ */
+#include <ctype.h>
+#include <dirent.h>
+#include <fcntl.h>
+#include <sys/mman.h>
+#include <sys/ptrace.h>
+#include <sys/resource.h>
+#include <sys/stat.h>
+#include <sys/syscall.h>
+#include <sys/wait.h>
+#include <unistd.h>
+#include <linux/io_uring.h>
+
+#include "coredump_test.h"
+
+#ifndef PTRACE_SETSIGMASK
+#define PTRACE_SETSIGMASK 0x420b
+#endif
+
+/* The dump of the tiny child takes well under a second. */
+#define EXIT_TIMEOUT_MS 5000
+
+FIXTURE_SETUP(coredump)
+{
+	FILE *file;
+	int ret;
+
+	self->pid_coredump_server = -ESRCH;
+	self->fd_tmpfs_detached = -1;
+	file = fopen("/proc/sys/kernel/core_pattern", "r");
+	ASSERT_NE(NULL, file);
+
+	ret = fread(self->original_core_pattern, 1, sizeof(self->original_core_pattern), file);
+	ASSERT_TRUE(ret || feof(file));
+	ASSERT_LT(ret, sizeof(self->original_core_pattern));
+
+	self->original_core_pattern[ret] = '\0';
+
+	ret = fclose(file);
+	ASSERT_EQ(0, ret);
+}
+
+FIXTURE_TEARDOWN(coredump)
+{
+	const char *reason;
+	FILE *file;
+	int ret;
+
+	file = fopen("/proc/sys/kernel/core_pattern", "w");
+	if (!file) {
+		reason = "Unable to open core_pattern";
+		goto fail;
+	}
+
+	ret = fprintf(file, "%s", self->original_core_pattern);
+	if (ret < 0) {
+		reason = "Unable to write to core_pattern";
+		goto fail;
+	}
+
+	ret = fclose(file);
+	if (ret) {
+		reason = "Unable to close core_pattern";
+		goto fail;
+	}
+
+	return;
+fail:
+	/* This should never happen */
+	fprintf(stderr, "Failed to cleanup coredump test: %s\n", reason);
+}
+
+/* A raw ring, no liburing. */
+struct uring {
+	int fd;
+	struct io_uring_params params;
+	void *sq;
+	size_t sq_len;
+	struct io_uring_sqe *sqes;
+	size_t sqes_len;
+	unsigned int *sq_tail, *sq_mask, *sq_array;
+	unsigned int *cq_head, *cq_tail, *cq_mask;
+	struct io_uring_cqe *cqes;
+};
+
+static int uring_setup(struct uring *r, unsigned int flags)
+{
+	size_t cq_len;
+
+	memset(r, 0, sizeof(*r));
+	r->params.flags = flags;
+	if (flags & IORING_SETUP_SQPOLL)
+		r->params.sq_thread_idle = 2000;
+	r->fd = syscall(__NR_io_uring_setup, 8, &r->params);
+	if (r->fd < 0)
+		return -1;
+	if (!(r->params.features & IORING_FEAT_SINGLE_MMAP))
+		return -1;
+
+	r->sq_len = r->params.sq_off.array + r->params.sq_entries * sizeof(unsigned int);
+	cq_len = r->params.cq_off.cqes + r->params.cq_entries * sizeof(struct io_uring_cqe);
+	if (cq_len > r->sq_len)
+		r->sq_len = cq_len;
+	r->sq = mmap(NULL, r->sq_len, PROT_READ | PROT_WRITE,
+		     MAP_SHARED | MAP_POPULATE, r->fd, IORING_OFF_SQ_RING);
+	if (r->sq == MAP_FAILED)
+		return -1;
+	r->sqes_len = r->params.sq_entries * sizeof(struct io_uring_sqe);
+	r->sqes = mmap(NULL, r->sqes_len, PROT_READ | PROT_WRITE,
+		       MAP_SHARED | MAP_POPULATE, r->fd, IORING_OFF_SQES);
+	if (r->sqes == MAP_FAILED)
+		return -1;
+
+	r->sq_tail = r->sq + r->params.sq_off.tail;
+	r->sq_mask = r->sq + r->params.sq_off.ring_mask;
+	r->sq_array = r->sq + r->params.sq_off.array;
+	r->cq_head = r->sq + r->params.cq_off.head;
+	r->cq_tail = r->sq + r->params.cq_off.tail;
+	r->cq_mask = r->sq + r->params.cq_off.ring_mask;
+	r->cqes = r->sq + r->params.cq_off.cqes;
+	return 0;
+}
+
+/* Submit one sqe, wait for its completion and return the result. */
+static int uring_submit_wait(struct uring *r, const struct io_uring_sqe *sqe)
+{
+	unsigned int tail = *r->sq_tail, idx = tail & *r->sq_mask;
+	unsigned int flags = IORING_ENTER_GETEVENTS;
+	int i;
+
+	r->sqes[idx] = *sqe;
+	r->sq_array[idx] = idx;
+	__atomic_store_n(r->sq_tail, tail + 1, __ATOMIC_RELEASE);
+
+	if (r->params.flags & IORING_SETUP_SQPOLL)
+		flags |= IORING_ENTER_SQ_WAKEUP;
+
+	for (i = 0; i < 100; i++) {
+		if (syscall(__NR_io_uring_enter, r->fd, 1, 1, flags, NULL, 0) < 0 &&
+		    errno != EINTR)
+			return -1;
+		if (__atomic_load_n(r->cq_tail, __ATOMIC_ACQUIRE) != *r->cq_head) {
+			unsigned int head = *r->cq_head;
+			int res = r->cqes[head & *r->cq_mask].res;
+
+			__atomic_store_n(r->cq_head, head + 1, __ATOMIC_RELEASE);
+			return res;
+		}
+		flags &= ~IORING_ENTER_SQ_WAKEUP;
+		usleep(10 * 1000);
+	}
+	return -1;
+}
+
+static bool uring_available(unsigned int flags)
+{
+	struct io_uring_params params = { .flags = flags };
+	int fd;
+
+	fd = syscall(__NR_io_uring_setup, 2, &params);
+	if (fd < 0)
+		return false;
+	close(fd);
+	return true;
+}
+
+/*
+ * Keep a ring, an idle io-wq worker and with SQPOLL an SQPOLL thread
+ * alive. The last worker of a ring never exits on its idle timeout.
+ */
+static void worker_child(bool sqpoll, int fd_ipc)
+{
+	struct rlimit rl = { RLIM_INFINITY, RLIM_INFINITY };
+	struct io_uring_sqe sqe = {};
+	static char buf[64];
+	struct uring ring;
+	int memfd;
+
+	if (setrlimit(RLIMIT_CORE, &rl))
+		_exit(EXIT_FAILURE);
+
+	memfd = memfd_create("coredump_worker", 0);
+	if (memfd < 0 || write(memfd, "hello", 5) != 5)
+		_exit(EXIT_FAILURE);
+
+	if (uring_setup(&ring, sqpoll ? IORING_SETUP_SQPOLL : 0))
+		_exit(EXIT_FAILURE);
+
+	/* IOSQE_ASYNC forces the read through io-wq so a worker appears. */
+	sqe.opcode = IORING_OP_READ;
+	sqe.fd = memfd;
+	sqe.addr = (__u64)(uintptr_t)buf;
+	sqe.len = sizeof(buf);
+	sqe.flags = IOSQE_ASYNC;
+	if (uring_submit_wait(&ring, &sqe) != 5)
+		_exit(EXIT_FAILURE);
+
+	if (write_nointr(fd_ipc, "1", 1) != 1)
+		_exit(EXIT_FAILURE);
+	close(fd_ipc);
+
+	for (;;)
+		pause();
+}
+
+/* Find the thread of @pid whose comm starts with @prefix. */
+static pid_t find_thread(pid_t pid, const char *prefix)
+{
+	char path[64], comm[64];
+	pid_t tid = -1;
+	struct dirent *de;
+	ssize_t bytes;
+	DIR *dir;
+	int fd;
+
+	snprintf(path, sizeof(path), "/proc/%d/task", pid);
+	dir = opendir(path);
+	if (!dir)
+		return -1;
+
+	while (tid < 0 && (de = readdir(dir))) {
+		if (!isdigit(de->d_name[0]))
+			continue;
+		snprintf(path, sizeof(path), "/proc/%d/task/%s/comm", pid, de->d_name);
+		fd = open(path, O_RDONLY | O_CLOEXEC);
+		if (fd < 0)
+			continue;
+		bytes = read(fd, comm, sizeof(comm) - 1);
+		close(fd);
+		if (bytes <= 0)
+			continue;
+		comm[bytes] = '\0';
+		if (!strncmp(comm, prefix, strlen(prefix)))
+			tid = atoi(de->d_name);
+	}
+	closedir(dir);
+	return tid;
+}
+
+/*
+ * Attach, stop the thread with SIGSTOP, drop the signal mask that
+ * copy_process() gave it and resume it with SIGSEGV instead.
+ */
+static bool inject_coredump_signal(pid_t pid, pid_t tid)
+{
+	__u64 mask = 0;
+	int status;
+
+	if (ptrace(PTRACE_SEIZE, tid, NULL, NULL))
+		return false;
+	if (syscall(SYS_tgkill, pid, tid, SIGSTOP))
+		return false;
+	if (waitpid(tid, &status, __WALL) != tid)
+		return false;
+	if (!WIFSTOPPED(status) || WSTOPSIG(status) != SIGSTOP)
+		return false;
+	if (ptrace(PTRACE_SETSIGMASK, tid, sizeof(mask), &mask))
+		return false;
+	return !ptrace(PTRACE_DETACH, tid, NULL, (void *)(long)SIGSEGV);
+}
+
+/* Reap @pid within @timeout_ms, -1 when it is still there. */
+static int wait_exit(pid_t pid, int *status, int timeout_ms)
+{
+	int i;
+
+	for (i = 0; i < timeout_ms / 10; i++) {
+		pid_t ret = waitpid(pid, status, WNOHANG);
+
+		if (ret == pid)
+			return 0;
+		if (ret < 0)
+			return -1;
+		usleep(10 * 1000);
+	}
+	return -1;
+}
+
+static void log_threads(struct __test_metadata *const _metadata, pid_t pid)
+{
+	char path[64], line[256], comm[64] = {};
+	struct dirent *de;
+	DIR *dir;
+	FILE *f;
+
+	snprintf(path, sizeof(path), "/proc/%d/task", pid);
+	dir = opendir(path);
+	if (!dir)
+		return;
+	while ((de = readdir(dir))) {
+		if (!isdigit(de->d_name[0]))
+			continue;
+		snprintf(path, sizeof(path), "/proc/%d/task/%s/status", pid, de->d_name);
+		f = fopen(path, "r");
+		if (!f)
+			continue;
+		while (fgets(line, sizeof(line), f)) {
+			line[strcspn(line, "\n")] = '\0';
+			if (!strncmp(line, "Name:", 5))
+				snprintf(comm, sizeof(comm), "%s", line + 6);
+			else if (!strncmp(line, "State:", 6))
+				TH_LOG("tid %s (%s) %s", de->d_name, comm, line + 7);
+		}
+		fclose(f);
+	}
+	closedir(dir);
+}
+
+enum dumper {
+	DUMPER_MAIN,
+	DUMPER_WORKER,
+	DUMPER_SQPOLL,
+};
+
+static void run_dumper(struct __test_metadata *const _metadata, bool sqpoll,
+		       enum dumper dumper)
+{
+	bool killed = false;
+	char path[64], c;
+	int ipc[2], status, fd;
+	pid_t pid, tid;
+
+	ASSERT_TRUE(set_core_pattern("/tmp/coredump.file.%p"));
+	ASSERT_EQ(pipe2(ipc, O_CLOEXEC), 0);
+
+	pid = fork();
+	ASSERT_GE(pid, 0);
+	if (pid == 0) {
+		close(ipc[0]);
+		worker_child(sqpoll, ipc[1]);
+	}
+	close(ipc[1]);
+	ASSERT_EQ(read_nointr(ipc[0], &c, 1), 1);
+	close(ipc[0]);
+
+	switch (dumper) {
+	case DUMPER_MAIN:
+		tid = pid;
+		break;
+	case DUMPER_WORKER:
+		tid = find_thread(pid, "iou-wrk-");
+		break;
+	case DUMPER_SQPOLL:
+		tid = find_thread(pid, "iou-sqp-");
+		break;
+	}
+	ASSERT_GT(tid, 0);
+	ASSERT_TRUE(inject_coredump_signal(pid, tid));
+
+	if (wait_exit(pid, &status, EXIT_TIMEOUT_MS)) {
+		/* No dump. Whatever happened, SIGKILL must still work. */
+		log_threads(_metadata, pid);
+		kill(pid, SIGKILL);
+		killed = true;
+		ASSERT_EQ(wait_exit(pid, &status, EXIT_TIMEOUT_MS), 0) {
+			TH_LOG("thread group %d is stuck after SIGSEGV to tid %d",
+			       pid, tid);
+		}
+	}
+
+	ASSERT_TRUE(WIFSIGNALED(status));
+	if (killed) {
+		TH_LOG("tid %d did not dump, the group was killed instead", tid);
+		ASSERT_EQ(WTERMSIG(status), SIGKILL);
+		return;
+	}
+	ASSERT_EQ(WTERMSIG(status), SIGSEGV);
+	ASSERT_TRUE(WCOREDUMP(status));
+
+	snprintf(path, sizeof(path), "/tmp/coredump.file.%d", pid);
+	fd = open(path, O_RDONLY | O_CLOEXEC);
+	unlink(path);
+	ASSERT_GE(fd, 0);
+	ASSERT_TRUE(check_coredump_extent(fd));
+	close(fd);
+}
+
+/* The mechanics: an injected SIGSEGV into a normal thread dumps core. */
+TEST_F(coredump, main_thread_dumper)
+{
+	if (!uring_available(0))
+		SKIP(return, "io_uring is not available");
+	run_dumper(_metadata, false, DUMPER_MAIN);
+}
+
+TEST_F(coredump, plain_worker_dumper)
+{
+	if (!uring_available(0))
+		SKIP(return, "io_uring is not available");
+	run_dumper(_metadata, false, DUMPER_WORKER);
+}
+
+TEST_F(coredump, sqpoll_thread_dumper)
+{
+	if (!uring_available(IORING_SETUP_SQPOLL))
+		SKIP(return, "io_uring SQPOLL is not available");
+	run_dumper(_metadata, true, DUMPER_SQPOLL);
+}
+
+/*
+ * The SQPOLL thread leaves its loop on the zap and waits for its io-wq
+ * workers to exit before it parks. The dumping worker never does.
+ */
+TEST_F(coredump, sqpoll_worker_dumper)
+{
+	if (!uring_available(IORING_SETUP_SQPOLL))
+		SKIP(return, "io_uring SQPOLL is not available");
+	run_dumper(_metadata, true, DUMPER_WORKER);
+}
+
+TEST_HARNESS_MAIN

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 11/17] selftests/coredump: expect PTRACE_SETSIGMASK to be refused on a user worker
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (9 preceding siblings ...)
  2026-09-21 13:44 ` [PATCH v3 10/17] selftests/coredump: test a user worker as the coredumping thread Christian Brauner
@ 2026-09-21 13:45 ` Christian Brauner
  2026-09-21 13:45 ` [PATCH v3 12/17] exec: cancel io_uring requests before de_thread() Christian Brauner
                   ` (5 subsequent siblings)
  16 siblings, 0 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:45 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable)

PTRACE_SETSIGMASK returns -EPERM for a user worker now.

- inject_coredump_signal() reports whether the mask could be changed
- a refused mask means the injected SIGSEGV stays pending on the worker
- the worker cases check that the group is alive and kill it

The main thread case still dumps.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 .../selftests/coredump/coredump_worker_test.c      | 44 ++++++++++++++++------
 1 file changed, 32 insertions(+), 12 deletions(-)

diff --git a/tools/testing/selftests/coredump/coredump_worker_test.c b/tools/testing/selftests/coredump/coredump_worker_test.c
index 81cadde0e752..9a4270b65a6e 100644
--- a/tools/testing/selftests/coredump/coredump_worker_test.c
+++ b/tools/testing/selftests/coredump/coredump_worker_test.c
@@ -16,6 +16,7 @@
  * SIGKILL still works. A failure leaves the stuck process behind.
  */
 #include <ctype.h>
+#include <errno.h>
 #include <dirent.h>
 #include <fcntl.h>
 #include <sys/mman.h>
@@ -255,24 +256,31 @@ static pid_t find_thread(pid_t pid, const char *prefix)
 
 /*
  * Attach, stop the thread with SIGSTOP, drop the signal mask that
- * copy_process() gave it and resume it with SIGSEGV instead.
+ * copy_process() gave it and resume it with SIGSEGV. Returns 1 when the
+ * mask was changed, 0 when PTRACE_SETSIGMASK was refused (a user worker
+ * keeps its mask and the SIGSEGV stays pending), -1 on any other failure.
  */
-static bool inject_coredump_signal(pid_t pid, pid_t tid)
+static int inject_coredump_signal(pid_t pid, pid_t tid)
 {
 	__u64 mask = 0;
-	int status;
+	int status, ret = 1;
 
 	if (ptrace(PTRACE_SEIZE, tid, NULL, NULL))
-		return false;
+		return -1;
 	if (syscall(SYS_tgkill, pid, tid, SIGSTOP))
-		return false;
+		return -1;
 	if (waitpid(tid, &status, __WALL) != tid)
-		return false;
+		return -1;
 	if (!WIFSTOPPED(status) || WSTOPSIG(status) != SIGSTOP)
-		return false;
-	if (ptrace(PTRACE_SETSIGMASK, tid, sizeof(mask), &mask))
-		return false;
-	return !ptrace(PTRACE_DETACH, tid, NULL, (void *)(long)SIGSEGV);
+		return -1;
+	if (ptrace(PTRACE_SETSIGMASK, tid, sizeof(mask), &mask)) {
+		if (errno != EPERM)
+			return -1;
+		ret = 0;
+	}
+	if (ptrace(PTRACE_DETACH, tid, NULL, (void *)(long)SIGSEGV))
+		return -1;
+	return ret;
 }
 
 /* Reap @pid within @timeout_ms, -1 when it is still there. */
@@ -333,7 +341,7 @@ static void run_dumper(struct __test_metadata *const _metadata, bool sqpoll,
 {
 	bool killed = false;
 	char path[64], c;
-	int ipc[2], status, fd;
+	int ipc[2], status, fd, ret;
 	pid_t pid, tid;
 
 	ASSERT_TRUE(set_core_pattern("/tmp/coredump.file.%p"));
@@ -361,7 +369,19 @@ static void run_dumper(struct __test_metadata *const _metadata, bool sqpoll,
 		break;
 	}
 	ASSERT_GT(tid, 0);
-	ASSERT_TRUE(inject_coredump_signal(pid, tid));
+	ret = inject_coredump_signal(pid, tid);
+	ASSERT_GE(ret, 0);
+	if (!ret) {
+		/* The signal sits on the worker, the group must be untouched. */
+		ASSERT_NE(dumper, DUMPER_MAIN);
+		TH_LOG("PTRACE_SETSIGMASK refused for tid %d, the SIGSEGV stays pending", tid);
+		ASSERT_EQ(wait_exit(pid, &status, 1000), -1);
+		kill(pid, SIGKILL);
+		ASSERT_EQ(wait_exit(pid, &status, EXIT_TIMEOUT_MS), 0);
+		ASSERT_TRUE(WIFSIGNALED(status));
+		ASSERT_EQ(WTERMSIG(status), SIGKILL);
+		return;
+	}
 
 	if (wait_exit(pid, &status, EXIT_TIMEOUT_MS)) {
 		/* No dump. Whatever happened, SIGKILL must still work. */

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 12/17] exec: cancel io_uring requests before de_thread()
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (10 preceding siblings ...)
  2026-09-21 13:45 ` [PATCH v3 11/17] selftests/coredump: expect PTRACE_SETSIGMASK to be refused on a user worker Christian Brauner
@ 2026-09-21 13:45 ` Christian Brauner
  2026-09-21 14:14   ` Oleg Nesterov
  2026-09-24 14:19   ` Jens Axboe
  2026-09-21 13:45 ` [PATCH v3 13/17] fork: move the coredump and exec checks into create_io_thread() Christian Brauner
                   ` (4 subsequent siblings)
  16 siblings, 2 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:45 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable),
	stable

begin_new_exec() cancels the exec'ing thread's io_uring requests after
de_thread() made it the only thread in the thread-group. Now
io_uring_task_cancel() runs task work after. If it runs a
create_worker_cb() item then a io-wq queued on the thread before the
exec creates a new io-wq worker thread.

Said thread joins the thread-group which de_thread() just took down. The
io-iwq sticks around until it notices that its workqueue is dying. But
code after de_thread() relies on being single-threaded.

Move io_uring_task_cancel() before de_thread() but behind the point of
no return. Any worker created during task work run from
io_uring_task_cancel() is just a sibling thread that de_thread() will
take down and reap.

Once io_uring_task_cancel() returned the thread has neither a task
context nor a workqueue left so nothing can add a thread behind
de_thread()'s back anymore.

Suggested-by: Oleg Nesterov <oleg@redhat.com>
Link: https://lore.kernel.org/r/aq0_lvHDViIpX_Mu@redhat.com
Fixes: 3bfe6106693b ("io-wq: fork worker threads from original task")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/exec.c | 11 +++++++----
 1 file changed, 7 insertions(+), 4 deletions(-)

diff --git a/fs/exec.c b/fs/exec.c
index 075a744421e1..334e6d358c56 100644
--- a/fs/exec.c
+++ b/fs/exec.c
@@ -1149,16 +1149,19 @@ int begin_new_exec(struct linux_binprm * bprm)
 	 */
 	bprm->point_of_no_return = true;
 
+	/*
+	 * Cancel any io_uring activity across execve. This runs task work
+	 * that may still create an io-wq worker, so do it while de_thread()
+	 * can still zap it.
+	 */
+	io_uring_task_cancel();
+
 	/* Make this the only thread in the thread group */
 	retval = de_thread(me);
 	if (retval)
 		goto out;
 	/* see the comment in check_unsafe_exec() */
 	current->fs->in_exec = 0;
-	/*
-	 * Cancel any io_uring activity across execve
-	 */
-	io_uring_task_cancel();
 
 	/* Ensure the files table is not shared. */
 	retval = unshare_fd(CLONE_FILES, &files);

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 13/17] fork: move the coredump and exec checks into create_io_thread()
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (11 preceding siblings ...)
  2026-09-21 13:45 ` [PATCH v3 12/17] exec: cancel io_uring requests before de_thread() Christian Brauner
@ 2026-09-21 13:45 ` Christian Brauner
  2026-09-21 14:15   ` Oleg Nesterov
  2026-09-21 13:45 ` [PATCH v3 14/17] fork: don't create io threads once PF_POSTCOREDUMP is set Christian Brauner
                   ` (3 subsequent siblings)
  16 siblings, 1 reply; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:45 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable)

Only a PF_USER_WORKER thread returns from get_signal() after a fatal
signal and can go on to create a thread. get_signal() sets PF_SIGNALED
before it dumps core and before it lets a PF_USER_WORKER thread return.
That flags sticks.

The exec'ing thread itself is never signaled. But it doesn't need the check
anyway as we moved cancellation of io_uring requests before de_thread().
create_io_thread() is the only way a signaled thread creates another
one. vhost_task_create() is reached from ioctls alone.

io_should_retry_thread() doesn't retry -EINTR, so io-wq callers give up
the same way they did when copy_process() refused.

Suggested-by: Oleg Nesterov <oleg@redhat.com>
Link: https://lore.kernel.org/r/aq0VLTyxNatFmLOK@redhat.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 kernel/fork.c | 4 ++++
 1 file changed, 4 insertions(+)

diff --git a/kernel/fork.c b/kernel/fork.c
index 50f5b3e2ca87..ede9f02bef47 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -2706,6 +2706,10 @@ struct task_struct *create_io_thread(int (*fn)(void *), void *arg, int node)
 		.user_worker	= 1,
 	};
 
+	/* A creator past its fatal signal gets no thread. */
+	if (current->flags & PF_SIGNALED)
+		return ERR_PTR(-EINTR);
+
 	return copy_process(NULL, 0, node, &args);
 }
 

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 14/17] fork: don't create io threads once PF_POSTCOREDUMP is set
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (12 preceding siblings ...)
  2026-09-21 13:45 ` [PATCH v3 13/17] fork: move the coredump and exec checks into create_io_thread() Christian Brauner
@ 2026-09-21 13:45 ` Christian Brauner
  2026-09-21 14:26   ` Oleg Nesterov
  2026-09-21 13:45 ` [PATCH v3 15/17] fork: use SIG_KERNEL_ONLY_MASK for the user worker signal mask Christian Brauner
                   ` (2 subsequent siblings)
  16 siblings, 1 reply; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:45 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable),
	stable

zap_process() skips every thread that already has PF_POSTCOREDUMP set.
Such a thread is past synchronize_group_exit(), so a coredump can't
catch it anymore. It isn't counted in core_state->threads_remaining and
it isn't sent SIGKILL.

do_exit() sets PF_POSTCOREDUMP in synchronize_group_exit() and calls
io_uring_files_cancel() right after that. The cancellation runs task
work and a create_worker_cb() that io-wq queued before the exit creates
a new io-wq worker from there. If another thread started a coredump in
the meantime that worker joins a thread-group which is already being
dumped. zap_process() never saw it so it was never counted. But when it
exits it sees signal->core_state in synchronize_group_exit() and
decrements threads_remaining like any other thread.

So the count reaches zero one thread early and the dumper leaves
coredump_wait_inactive() while a thread it counted is still running.

The task has no fatal signal pending because zap_process() deliberately
didn't send it one, and it never went through get_signal() so it doesn't
have PF_SIGNALED either.

Refuse to create an io thread when the creator has PF_POSTCOREDUMP.
That costs nothing. io_uring_files_cancel() raises IO_WQ_BIT_EXIT
before it runs any task work, so a worker created from there only ever
gets to exit again, and io_should_retry_thread() doesn't retry -EINTR.

Clearing PF_POSTCOREDUMP for the new thread in copy_process() was the
other option. It makes the worker a thread like any other, but it only
helps while the dump hasn't started. A worker born after zap_process()
has run is invisible to it whatever its flags say and still decrements
threads_remaining on the way out. Refusing to create it covers both, and
then no task is ever born with the flag, so there is nothing to clear.

Moving io_uring_files_cancel() ahead of synchronize_group_exit() closes
the same window. It runs the cancellation, and whatever that can block
on, before the thread announces itself to the dumper. A thread that
blocks there never reaches coredump_task_exit() at all, so the dumper
ends up waiting for a thread that never parks. Refusing the creation
leaves the cancellation where it is. But it's very ugly to run io_uring
work even before we did all the generic exit work.

Fixes: 92307383082d ("coredump:  Don't perform any cleanups before dumping core")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 kernel/fork.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/kernel/fork.c b/kernel/fork.c
index ede9f02bef47..fd2829b0a8b6 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -2706,8 +2706,8 @@ struct task_struct *create_io_thread(int (*fn)(void *), void *arg, int node)
 		.user_worker	= 1,
 	};
 
-	/* A creator past its fatal signal gets no thread. */
-	if (current->flags & PF_SIGNALED)
+	/* A creator past its fatal signal or its coredump point gets no thread. */
+	if (current->flags & (PF_SIGNALED | PF_POSTCOREDUMP))
 		return ERR_PTR(-EINTR);
 
 	return copy_process(NULL, 0, node, &args);

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 15/17] fork: use SIG_KERNEL_ONLY_MASK for the user worker signal mask
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (13 preceding siblings ...)
  2026-09-21 13:45 ` [PATCH v3 14/17] fork: don't create io threads once PF_POSTCOREDUMP is set Christian Brauner
@ 2026-09-21 13:45 ` Christian Brauner
  2026-09-21 14:29   ` Oleg Nesterov
  2026-09-21 13:45 ` [PATCH v3 16/17] signal: enforce the user worker signal mask in __set_task_blocked() Christian Brauner
  2026-09-21 13:45 ` [PATCH v3 17/17] fs: close files from the highest descriptor down Christian Brauner
  16 siblings, 1 reply; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:45 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable)

For a user worker copy_process() blocks every signal except for SIGKILL
and SIGSTOP. Use SIG_KERNEL_ONLY_MASK instead of open-coding it.

No functional changes.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 kernel/fork.c | 7 ++-----
 1 file changed, 2 insertions(+), 5 deletions(-)

diff --git a/kernel/fork.c b/kernel/fork.c
index fd2829b0a8b6..be5980d0dd9f 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -2146,12 +2146,9 @@ __latent_entropy struct task_struct *copy_process(
 	if (args->kthread)
 		p->flags |= PF_KTHREAD;
 	if (args->user_worker) {
-		/*
-		 * Mark us a user worker, and block any signal that isn't
-		 * fatal or STOP
-		 */
+		/* A user worker takes only the signals nobody can block. */
 		p->flags |= PF_USER_WORKER;
-		siginitsetinv(&p->blocked, sigmask(SIGKILL)|sigmask(SIGSTOP));
+		siginitsetinv(&p->blocked, SIG_KERNEL_ONLY_MASK);
 	}
 	if (args->io_thread)
 		p->flags |= PF_IO_WORKER;

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 16/17] signal: enforce the user worker signal mask in __set_task_blocked()
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (14 preceding siblings ...)
  2026-09-21 13:45 ` [PATCH v3 15/17] fork: use SIG_KERNEL_ONLY_MASK for the user worker signal mask Christian Brauner
@ 2026-09-21 13:45 ` Christian Brauner
  2026-09-21 16:16   ` Oleg Nesterov
  2026-09-21 13:45 ` [PATCH v3 17/17] fs: close files from the highest descriptor down Christian Brauner
  16 siblings, 1 reply; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:45 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable)

User workers block every signal except for SIGKILL and SIGSTOP in
copy_process(). So complete_signal() never picks a thread that blocks
the signal and get_signal() never dequeues such a thread. Net effect is
that user workers only exit on SIGKILL or stop on SIGSTOP.

Before we fixed it a tracer could unblock a signal. For example, by
unblocking SIGHUP a installing a handler complete_signal() could end up
picking a user worker for a process-directed SIGHUP. Which effectively
means the handler never runs.

After blocking PTRACE_SETSIGMASK what remains is the in-kernel
sigprocmask() and force_sig_info_to_task() changes.

The in-kernel sigprocmask() users block more signals for a critical
section and restore the saved mask afterwards. cifs used to do that in
__smb_send_rqst() and ocfs2 in ocfs2_block_signals(). Both were
reachable from an io-wq worker. Both only add SIGKILL and SIGSTOP to the
user workers's and then put the fork-time mask back.

force_sig_info_to_task() unblocks a signal so a target can't hide from
the signal. A synchronous signal forced onto a user worker is accepted
currently so let's leave that alone.

Make it a rule that a user worker's signal mask can never drop below
the copy_process() deafult. SIGKILL and SIGSTOP have the same rule in
the other direction. rt_sigprocmask(), set_current_blocked() and
PTRACE_SETSIGMASK strip them from whatever mask userspace asks for.

Do the same for user worker mask in __set_task_blocked(). Add the
fork-time mask back into the new set for a user worker and warn if that
changed anything. Warn when a caller unblocks a signal for a user worker.

No functional changes.

Link: https://lore.kernel.org/r/aq_4fY6GVtK47Njq@redhat.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 kernel/signal.c | 10 ++++++++++
 1 file changed, 10 insertions(+)

diff --git a/kernel/signal.c b/kernel/signal.c
index c3bad983dd37..137002445b0d 100644
--- a/kernel/signal.c
+++ b/kernel/signal.c
@@ -3212,6 +3212,16 @@ long do_no_restart_syscall(struct restart_block *param)
 
 static void __set_task_blocked(struct task_struct *tsk, const sigset_t *newset)
 {
+	sigset_t floor, floored;
+
+	/* A user worker never unblocks anything but SIGKILL and SIGSTOP. */
+	if (unlikely(tsk->flags & PF_USER_WORKER)) {
+		siginitsetinv(&floor, SIG_KERNEL_ONLY_MASK);
+		sigorsets(&floored, newset, &floor);
+		WARN_ON_ONCE(!sigequalsets(&floored, newset));
+		newset = &floored;
+	}
+
 	if (task_sigpending(tsk) && !thread_group_empty(tsk)) {
 		sigset_t newblocked;
 		/* A set of now blocked but previously unblocked signals. */

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v3 17/17] fs: close files from the highest descriptor down
  2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
                   ` (15 preceding siblings ...)
  2026-09-21 13:45 ` [PATCH v3 16/17] signal: enforce the user worker signal mask in __set_task_blocked() Christian Brauner
@ 2026-09-21 13:45 ` Christian Brauner
  16 siblings, 0 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 13:45 UTC (permalink / raw)
  To: Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Jens Axboe, Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar,
	Peter Zijlstra, linux-mm, io-uring, Christian Brauner (Amutable)

close_files(), __range_close() and close_cloexec_files() walk the
descriptor table from the lowest descriptor up. They used to call
filp_close(), which left the final __fput() to task work. task work is a
LIFO list, so the releases ran after the walk had finished and in the
opposite direction, highest descriptor first.

This dumb ordering is relevant for a bunch of broken but long-standing
cases. It matters whenever the ->flush() or ->release() of one file
waits for something that only the release of another file of the same
table provides. Then one of the two orders deadlocks and the other one
doesn't:

    exit, fd 3 is one end of a pipe    peer
    -------------------------------    ----
                                       splice(socket -> pipe)
                                         pipe_lock()
                                         waits for data or EOF
    close_files()
      fd 3: pipe_release()
        mutex_lock(&pipe->mutex)         held by the peer
      fd 5: the socket, not reached      would be the peer's EOF

Programs create the thing that guards or wakes another thing first.
Hence, it gets the lower descriptor which is the layout that breaks:

(1) a tap device released before the AF_LLC socket that holds a
    reference to it
(2) an unlinked fsdax file evicted before the pipe that holds its
    vmspliced pages
(3) an overlayfs directory whose release queues up behind an unlink that
    waits for a splice into the same directory

The exiting task is unkillable in all of them. All of that crap can
obviously also become a bug if you reorder the file descriptors.

Continue walking all three tables from the highest descriptor down. That
restores the order the deferred puts had. ->flush() moves with the
release. So it now runs highest descriptor first as well.

None of this fixes the underlying defects. For every one of these pairs
the mirrored layout deadlocked before and deadlocks
again now:

(1') splice() holding pipe->mutex across unbounded socket and tty I/O
(2') AF_LLC keeping a netdev reference without a NETDEV_UNREGISTER handler
(3') uninterruptible wait in dax_break_layout_final()
(4') ovl_splice_write() sleeping under the inode lock

It all predates the synchronous close and each should really get fixed.

Reported-by: Chris Mason <mason@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/file.c | 55 +++++++++++++++++++++++++++++--------------------------
 1 file changed, 29 insertions(+), 26 deletions(-)

diff --git a/fs/file.c b/fs/file.c
index 76e328edf630..7f8d0afd8807 100644
--- a/fs/file.c
+++ b/fs/file.c
@@ -497,24 +497,21 @@ static struct fdtable *close_files(struct files_struct *files)
 	 * files structure.
 	 */
 	struct fdtable *fdt = rcu_dereference_raw(files->fdt);
-	unsigned int i, j = 0;
+	unsigned int j = fdt->max_fds / BITS_PER_LONG;
+
+	/* Highest fd first, the order the deferred puts ran in. */
+	while (j--) {
+		unsigned long set = fdt->open_fds[j];
 
-	for (;;) {
-		unsigned long set;
-		i = j * BITS_PER_LONG;
-		if (i >= fdt->max_fds)
-			break;
-		set = fdt->open_fds[j++];
 		while (set) {
-			if (set & 1) {
-				struct file *file = fdt->fd[i];
-				if (file) {
-					filp_close_sync(file, files);
-					cond_resched();
-				}
+			unsigned int bit = __fls(set);
+			struct file *file = fdt->fd[j * BITS_PER_LONG + bit];
+
+			set ^= 1UL << bit;
+			if (file) {
+				filp_close_sync(file, files);
+				cond_resched();
 			}
-			i++;
-			set >>= 1;
 		}
 	}
 
@@ -801,10 +798,14 @@ static inline void __range_close(struct files_struct *files, unsigned int fd,
 	n = last_fd(fdt);
 	max_fd = min(max_fd, n);
 
-	for (fd = find_next_bit(fdt->open_fds, max_fd + 1, fd);
-	     fd <= max_fd;
-	     fd = find_next_bit(fdt->open_fds, max_fd + 1, fd + 1)) {
-		file = file_close_fd_locked(files, fd);
+	/* Highest fd first, see close_files(). */
+	for (n = max_fd + 1; n > fd; ) {
+		unsigned int cur = find_last_bit(fdt->open_fds, n);
+
+		if (cur >= n || cur < fd)
+			break;
+		n = cur;
+		file = file_close_fd_locked(files, cur);
 		if (file) {
 			spin_unlock(&files->file_lock);
 			filp_close_sync(file, files);
@@ -908,20 +909,22 @@ void close_cloexec_files(struct files_struct *files)
 
 	/* exec unshares first */
 	spin_lock(&files->file_lock);
-	for (i = 0; ; i++) {
+	fdt = files_fdtable(files);
+	/* Highest fd first, see close_files(). */
+	for (i = fdt->max_fds / BITS_PER_LONG; i--; ) {
 		unsigned long set;
-		unsigned fd = i * BITS_PER_LONG;
+
 		fdt = files_fdtable(files);
-		if (fd >= fdt->max_fds)
-			break;
 		set = fdt->close_on_exec[i];
 		if (!set)
 			continue;
 		fdt->close_on_exec[i] = 0;
-		for ( ; set ; fd++, set >>= 1) {
+		while (set) {
+			unsigned int bit = __fls(set);
+			unsigned fd = i * BITS_PER_LONG + bit;
 			struct file *file;
-			if (!(set & 1))
-				continue;
+
+			set ^= 1UL << bit;
 			file = fdt->fd[fd];
 			if (!file)
 				continue;

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* Re: [PATCH v3 09/17] ptrace: refuse to change the signal mask of a user worker
  2026-09-21 13:44 ` [PATCH v3 09/17] ptrace: refuse to change the signal mask of a user worker Christian Brauner
@ 2026-09-21 14:14   ` Oleg Nesterov
  0 siblings, 0 replies; 29+ messages in thread
From: Oleg Nesterov @ 2026-09-21 14:14 UTC (permalink / raw)
  To: Christian Brauner
  Cc: Chris Mason, linux-fsdevel, Jens Axboe, Alexander Viro, Jan Kara,
	NeilBrown, Ingo Molnar, Peter Zijlstra, linux-mm, io-uring,
	stable

On 09/21, Christian Brauner wrote:
>
> --- a/kernel/ptrace.c
> +++ b/kernel/ptrace.c
> @@ -1227,6 +1227,12 @@ int ptrace_request(struct task_struct *child, long request,
>  	case PTRACE_SETSIGMASK: {
>  		sigset_t new_set;
>
> +		/* A user worker only ever takes SIGKILL and SIGSTOP. */
> +		if (child->flags & PF_USER_WORKER) {
> +			ret = -EPERM;
> +			break;
> +		}

Acked-by: Oleg Nesterov <oleg@redhat.com>


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v3 12/17] exec: cancel io_uring requests before de_thread()
  2026-09-21 13:45 ` [PATCH v3 12/17] exec: cancel io_uring requests before de_thread() Christian Brauner
@ 2026-09-21 14:14   ` Oleg Nesterov
  2026-09-24 14:19   ` Jens Axboe
  1 sibling, 0 replies; 29+ messages in thread
From: Oleg Nesterov @ 2026-09-21 14:14 UTC (permalink / raw)
  To: Christian Brauner
  Cc: Chris Mason, linux-fsdevel, Jens Axboe, Alexander Viro, Jan Kara,
	NeilBrown, Ingo Molnar, Peter Zijlstra, linux-mm, io-uring,
	stable

On 09/21, Christian Brauner wrote:
>
> --- a/fs/exec.c
> +++ b/fs/exec.c
> @@ -1149,16 +1149,19 @@ int begin_new_exec(struct linux_binprm * bprm)
>  	 */
>  	bprm->point_of_no_return = true;
>  
> +	/*
> +	 * Cancel any io_uring activity across execve. This runs task work
> +	 * that may still create an io-wq worker, so do it while de_thread()
> +	 * can still zap it.
> +	 */
> +	io_uring_task_cancel();

Acked-by: Oleg Nesterov <oleg@redhat.com>


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v3 13/17] fork: move the coredump and exec checks into create_io_thread()
  2026-09-21 13:45 ` [PATCH v3 13/17] fork: move the coredump and exec checks into create_io_thread() Christian Brauner
@ 2026-09-21 14:15   ` Oleg Nesterov
  0 siblings, 0 replies; 29+ messages in thread
From: Oleg Nesterov @ 2026-09-21 14:15 UTC (permalink / raw)
  To: Christian Brauner
  Cc: Chris Mason, linux-fsdevel, Jens Axboe, Alexander Viro, Jan Kara,
	NeilBrown, Ingo Molnar, Peter Zijlstra, linux-mm, io-uring

On 09/21, Christian Brauner wrote:
>
> --- a/kernel/fork.c
> +++ b/kernel/fork.c
> @@ -2706,6 +2706,10 @@ struct task_struct *create_io_thread(int (*fn)(void *), void *arg, int node)
>  		.user_worker	= 1,
>  	};
>  
> +	/* A creator past its fatal signal gets no thread. */
> +	if (current->flags & PF_SIGNALED)
> +		return ERR_PTR(-EINTR);

Acked-by: Oleg Nesterov <oleg@redhat.com>


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v3 14/17] fork: don't create io threads once PF_POSTCOREDUMP is set
  2026-09-21 13:45 ` [PATCH v3 14/17] fork: don't create io threads once PF_POSTCOREDUMP is set Christian Brauner
@ 2026-09-21 14:26   ` Oleg Nesterov
  0 siblings, 0 replies; 29+ messages in thread
From: Oleg Nesterov @ 2026-09-21 14:26 UTC (permalink / raw)
  To: Christian Brauner
  Cc: Chris Mason, linux-fsdevel, Jens Axboe, Alexander Viro, Jan Kara,
	NeilBrown, Ingo Molnar, Peter Zijlstra, linux-mm, io-uring,
	stable

On 09/21, Christian Brauner wrote:
>
> Refuse to create an io thread when the creator has PF_POSTCOREDUMP.
> That costs nothing. io_uring_files_cancel() raises IO_WQ_BIT_EXIT
> before it runs any task work, so a worker created from there only ever
> gets to exit again, and io_should_retry_thread() doesn't retry -EINTR.

Thanks for this note in the changelog,

> --- a/kernel/fork.c
> +++ b/kernel/fork.c
> @@ -2706,8 +2706,8 @@ struct task_struct *create_io_thread(int (*fn)(void *), void *arg, int node)
>  		.user_worker	= 1,
>  	};
>
> -	/* A creator past its fatal signal gets no thread. */
> -	if (current->flags & PF_SIGNALED)
> +	/* A creator past its fatal signal or its coredump point gets no thread. */
> +	if (current->flags & (PF_SIGNALED | PF_POSTCOREDUMP))
>  		return ERR_PTR(-EINTR);

Acked-by: Oleg Nesterov <oleg@redhat.com>


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v3 15/17] fork: use SIG_KERNEL_ONLY_MASK for the user worker signal mask
  2026-09-21 13:45 ` [PATCH v3 15/17] fork: use SIG_KERNEL_ONLY_MASK for the user worker signal mask Christian Brauner
@ 2026-09-21 14:29   ` Oleg Nesterov
  0 siblings, 0 replies; 29+ messages in thread
From: Oleg Nesterov @ 2026-09-21 14:29 UTC (permalink / raw)
  To: Christian Brauner
  Cc: Chris Mason, linux-fsdevel, Jens Axboe, Alexander Viro, Jan Kara,
	NeilBrown, Ingo Molnar, Peter Zijlstra, linux-mm, io-uring

On 09/21, Christian Brauner wrote:
>
> --- a/kernel/fork.c
> +++ b/kernel/fork.c
> @@ -2146,12 +2146,9 @@ __latent_entropy struct task_struct *copy_process(
>  	if (args->kthread)
>  		p->flags |= PF_KTHREAD;
>  	if (args->user_worker) {
> -		/*
> -		 * Mark us a user worker, and block any signal that isn't
> -		 * fatal or STOP
> -		 */
> +		/* A user worker takes only the signals nobody can block. */
>  		p->flags |= PF_USER_WORKER;
> -		siginitsetinv(&p->blocked, sigmask(SIGKILL)|sigmask(SIGSTOP));
> +		siginitsetinv(&p->blocked, SIG_KERNEL_ONLY_MASK);

Acked-by: Oleg Nesterov <oleg@redhat.com>


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v3 16/17] signal: enforce the user worker signal mask in __set_task_blocked()
  2026-09-21 13:45 ` [PATCH v3 16/17] signal: enforce the user worker signal mask in __set_task_blocked() Christian Brauner
@ 2026-09-21 16:16   ` Oleg Nesterov
  2026-09-21 20:05     ` Christian Brauner
  0 siblings, 1 reply; 29+ messages in thread
From: Oleg Nesterov @ 2026-09-21 16:16 UTC (permalink / raw)
  To: Christian Brauner
  Cc: Chris Mason, linux-fsdevel, Jens Axboe, Alexander Viro, Jan Kara,
	NeilBrown, Ingo Molnar, Peter Zijlstra, linux-mm, io-uring

On 09/21, Christian Brauner wrote:
>
> The in-kernel sigprocmask() users block more signals for a critical
> section and restore the saved mask afterwards. cifs used to do that in
> __smb_send_rqst() and ocfs2 in ocfs2_block_signals(). Both were
> reachable from an io-wq worker. Both only add SIGKILL and SIGSTOP to the
> user workers's and then put the fork-time mask back.

Ooh, this reminds me... see below.

> --- a/kernel/signal.c
> +++ b/kernel/signal.c
> @@ -3212,6 +3212,16 @@ long do_no_restart_syscall(struct restart_block *param)
>
>  static void __set_task_blocked(struct task_struct *tsk, const sigset_t *newset)
>  {
> +	sigset_t floor, floored;
> +
> +	/* A user worker never unblocks anything but SIGKILL and SIGSTOP. */
> +	if (unlikely(tsk->flags & PF_USER_WORKER)) {
> +		siginitsetinv(&floor, SIG_KERNEL_ONLY_MASK);
> +		sigorsets(&floored, newset, &floor);
> +		WARN_ON_ONCE(!sigequalsets(&floored, newset));
> +		newset = &floored;
> +	}

OK, Acked-by: Oleg Nesterov <oleg@redhat.com>

although to me something like a more simple version

	/* A user worker never unblocks anything but SIGKILL and SIGSTOP. */
	if (unlikely(tsk->flags & PF_USER_WORKER)) {
		sigset_t xxx;
		siginitsetinv(&xxx, SIG_KERNEL_ONLY_MASK);
		if (WARN_ON_ONCE(has_pending_signals(&xxx, newset)))
			return;
	}

makes more sense. But this is minor.

In fact I think that __set_task_blocked() should simply do

	if (WARN_ON_ONCE(tsk->flags & PF_USER_WORKER))
		return;

but yes, we can't do this right now, we have in-kernel abusers of sigprocmask().

IMO, they should be changed to not rely on sigprocmask(). Lets look at
ocfs2_delete_inode() for example,

	/* We want to block signals in delete_inode as the lock and
	 * messaging paths may return us -ERESTARTSYS. Which would
	 * cause us to exit early, resulting in inodes being orphaned
	 * forever. */
	ocfs2_block_signals(&oldset);

ocfs2_block_signals() blocks everything including SIGKILL and SIGSTOP.
I guess to protect against signal_wake_up / TIF_SIGPENDING ?

But this can only help in the single-threaded case. And PF_USER_WORKER's
are never single-threaded.

Blocking SIGSTOP can't protect from SIGSTOP if another thread dequeues
SIGSTOP or another sig_kernel_stop() signal. This another thread will do
do_signal_stop() -> signal_wake_up().

Same for SIGKILL...

In short, I agree with this patch, but mostly because of WARN_ON_ONCE()
it adds.

Oleg.


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v3 16/17] signal: enforce the user worker signal mask in __set_task_blocked()
  2026-09-21 16:16   ` Oleg Nesterov
@ 2026-09-21 20:05     ` Christian Brauner
  0 siblings, 0 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-21 20:05 UTC (permalink / raw)
  To: Oleg Nesterov
  Cc: Chris Mason, linux-fsdevel, Jens Axboe, Alexander Viro, Jan Kara,
	NeilBrown, Ingo Molnar, Peter Zijlstra, linux-mm, io-uring

On Mon, Sep 21, 2026 at 06:16:24PM +0200, Oleg Nesterov wrote:
> On 09/21, Christian Brauner wrote:
> >
> > The in-kernel sigprocmask() users block more signals for a critical
> > section and restore the saved mask afterwards. cifs used to do that in
> > __smb_send_rqst() and ocfs2 in ocfs2_block_signals(). Both were
> > reachable from an io-wq worker. Both only add SIGKILL and SIGSTOP to the
> > user workers's and then put the fork-time mask back.
> 
> Ooh, this reminds me... see below.
> 
> > --- a/kernel/signal.c
> > +++ b/kernel/signal.c
> > @@ -3212,6 +3212,16 @@ long do_no_restart_syscall(struct restart_block *param)
> >
> >  static void __set_task_blocked(struct task_struct *tsk, const sigset_t *newset)
> >  {
> > +	sigset_t floor, floored;
> > +
> > +	/* A user worker never unblocks anything but SIGKILL and SIGSTOP. */
> > +	if (unlikely(tsk->flags & PF_USER_WORKER)) {
> > +		siginitsetinv(&floor, SIG_KERNEL_ONLY_MASK);
> > +		sigorsets(&floored, newset, &floor);
> > +		WARN_ON_ONCE(!sigequalsets(&floored, newset));
> > +		newset = &floored;
> > +	}
> 
> OK, Acked-by: Oleg Nesterov <oleg@redhat.com>
> 
> although to me something like a more simple version
> 
> 	/* A user worker never unblocks anything but SIGKILL and SIGSTOP. */
> 	if (unlikely(tsk->flags & PF_USER_WORKER)) {
> 		sigset_t xxx;
> 		siginitsetinv(&xxx, SIG_KERNEL_ONLY_MASK);
> 		if (WARN_ON_ONCE(has_pending_signals(&xxx, newset)))
> 			return;
> 	}

Ok, let me see and fold that in if this is workable.

> makes more sense. But this is minor.
> 
> In fact I think that __set_task_blocked() should simply do
> 
> 	if (WARN_ON_ONCE(tsk->flags & PF_USER_WORKER))
> 		return;
> 
> but yes, we can't do this right now, we have in-kernel abusers of sigprocmask().

Yes, I found the two offenders. But fwiw, I've already fixed smb with
PF_NO_NOTIFY_SIGNAL which you reviewed some time ago. It's sitting in
kernel-7.4.signal. So only ocfs2 is left.

> IMO, they should be changed to not rely on sigprocmask(). Lets look at
> ocfs2_delete_inode() for example,
> 
> 	/* We want to block signals in delete_inode as the lock and
> 	 * messaging paths may return us -ERESTARTSYS. Which would
> 	 * cause us to exit early, resulting in inodes being orphaned
> 	 * forever. */
> 	ocfs2_block_signals(&oldset);
> 
> ocfs2_block_signals() blocks everything including SIGKILL and SIGSTOP.
> I guess to protect against signal_wake_up / TIF_SIGPENDING ?

I don't know. I guess so.

> But this can only help in the single-threaded case. And PF_USER_WORKER's
> are never single-threaded.

Yes.

> Blocking SIGSTOP can't protect from SIGSTOP if another thread dequeues
> SIGSTOP or another sig_kernel_stop() signal. This another thread will do
> do_signal_stop() -> signal_wake_up().
> 
> Same for SIGKILL...

Yes.

> In short, I agree with this patch, but mostly because of WARN_ON_ONCE()
> it adds.

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v3 02/17] signal: only SIGKILL and the freezers interrupt a coredumping task
  2026-09-21 13:44 ` [PATCH v3 02/17] signal: only SIGKILL and the freezers interrupt a coredumping task Christian Brauner
@ 2026-09-22 12:55   ` Oleg Nesterov
  2026-09-22 14:33     ` Christian Brauner
  0 siblings, 1 reply; 29+ messages in thread
From: Oleg Nesterov @ 2026-09-22 12:55 UTC (permalink / raw)
  To: Christian Brauner
  Cc: Chris Mason, linux-fsdevel, Jens Axboe, Alexander Viro, Jan Kara,
	NeilBrown, Ingo Molnar, Peter Zijlstra, linux-mm, io-uring,
	stable

On 09/21, Christian Brauner wrote:
>
> For quite a while now coredump has allowed either SIGKILL or freezing to
> interrupt an ongoing coredump. Since we seem to like it complicated
> "freezing" can mean a lot of things:
>
> (1) Power Management induced freezing
> (2) cgroup v1 freezer controller induced freezing
> (3) cgroup v2 freezer controller induced freezing

All of them do signal_wake_up() which sets TIF_SIGPENDING
However, zap_threads() clears this flag.

With the recent changes including your PF_NO_NOTIFY_SIGNAL and
05/17 "signal: don't retarget shared signals in a dying thread group",
what else can make signal_pending() true? Apart from SIGKILL.

So perhaps we can simply do _something_ like

	--- x/fs/coredump.c
	+++ x/fs/coredump.c
	@@ -508,7 +508,8 @@ static int zap_threads(struct task_struc
		int nr = -EAGAIN;
	 
		spin_lock_irq(&tsk->sighand->siglock);
	-	if (!(signal->flags & SIGNAL_GROUP_EXIT) && !signal->group_exec_task) {
	+	if (!(signal->flags & SIGNAL_GROUP_EXIT) && !signal->group_exec_task
	+	    && !freezing(tsk) && !(tsk->jobctl & JOBCTL_TRAP_FREEZE)) {
			/* Allow SIGKILL, see prepare_signal() */
			signal->core_state = core_state;
			nr = zap_process(signal, exit_code);
	@@ -583,7 +584,7 @@ static bool dump_interrupted(void)
		 * but then we need to teach dump_write() to restart and clear
		 * TIF_SIGPENDING.
		 */
	-	return fatal_signal_pending(current) || freezing(current);
	+	return signal_pending(current);
	 }
	 
	 static void wait_for_dump_helpers(struct file *file)

?

At first glance the change in dump_interrupted() makes sense anyway.
For example, dumping to pipe can't tolerate signal_pendin() == true.

What do you think?

> Fixes: 403bad72b67d ("coredump: only SIGKILL should interrupt the coredumping task")

Well, I think we should blame 76f969e8948d ("cgroup: cgroup v2 freezer")

Oleg.


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v3 02/17] signal: only SIGKILL and the freezers interrupt a coredumping task
  2026-09-22 12:55   ` Oleg Nesterov
@ 2026-09-22 14:33     ` Christian Brauner
  0 siblings, 0 replies; 29+ messages in thread
From: Christian Brauner @ 2026-09-22 14:33 UTC (permalink / raw)
  To: Oleg Nesterov
  Cc: Chris Mason, linux-fsdevel, Jens Axboe, Alexander Viro, Jan Kara,
	NeilBrown, Ingo Molnar, Peter Zijlstra, linux-mm, io-uring,
	stable

On Tue, Sep 22, 2026 at 02:55:41PM +0200, Oleg Nesterov wrote:
> On 09/21, Christian Brauner wrote:
> >
> > For quite a while now coredump has allowed either SIGKILL or freezing to
> > interrupt an ongoing coredump. Since we seem to like it complicated
> > "freezing" can mean a lot of things:
> >
> > (1) Power Management induced freezing
> > (2) cgroup v1 freezer controller induced freezing
> > (3) cgroup v2 freezer controller induced freezing
> 
> All of them do signal_wake_up() which sets TIF_SIGPENDING
> However, zap_threads() clears this flag.
> 
> With the recent changes including your PF_NO_NOTIFY_SIGNAL and
> 05/17 "signal: don't retarget shared signals in a dying thread group",
> what else can make signal_pending() true? Apart from SIGKILL.
> 
> So perhaps we can simply do _something_ like
> 
> 	--- x/fs/coredump.c
> 	+++ x/fs/coredump.c
> 	@@ -508,7 +508,8 @@ static int zap_threads(struct task_struc
> 		int nr = -EAGAIN;
> 	 
> 		spin_lock_irq(&tsk->sighand->siglock);
> 	-	if (!(signal->flags & SIGNAL_GROUP_EXIT) && !signal->group_exec_task) {
> 	+	if (!(signal->flags & SIGNAL_GROUP_EXIT) && !signal->group_exec_task
> 	+	    && !freezing(tsk) && !(tsk->jobctl & JOBCTL_TRAP_FREEZE)) {
> 			/* Allow SIGKILL, see prepare_signal() */
> 			signal->core_state = core_state;
> 			nr = zap_process(signal, exit_code);
> 	@@ -583,7 +584,7 @@ static bool dump_interrupted(void)
> 		 * but then we need to teach dump_write() to restart and clear
> 		 * TIF_SIGPENDING.
> 		 */
> 	-	return fatal_signal_pending(current) || freezing(current);
> 	+	return signal_pending(current);
> 	 }
> 	 
> 	 static void wait_for_dump_helpers(struct file *file)
> 
> ?
> 
> At first glance the change in dump_interrupted() makes sense anyway.
> For example, dumping to pipe can't tolerate signal_pendin() == true.
> 
> What do you think?

Yes. Here's what I have in the tree.

Subject: [PATCH v4 02/17] signal: only SIGKILL and the freezers interrupt a
 coredumping task

For quite a while now coredump has allowed either SIGKILL or freezing to
interrupt an ongoing coredump. Since we seem to like it complicated
"freezing" can mean a lot of things:

(1) Power Management induced freezing
(2) cgroup v1 freezer controller induced freezing
(3) cgroup v2 freezer controller induced freezing

So (1) and (2) are handled by checking freezing() but (3) isn't.
(3) uses cgroup_freeze_task() which sets JOBCTL_TRAP_FREEZE and calls
signal_wake_up() on every task in the cgroup. The trap isn't handled by
the coredump code.

Fixes: 76f969e8948d ("cgroup: cgroup v2 freezer")
Cc: stable@vger.kernel.org
Suggested-by: Oleg Nesterov <oleg@redhat.com>
Link: https://lore.kernel.org/r/arJ6zZuI59BJnXpl@redhat.com
Link: https://patch.msgid.link/20260921-work-coredump-fixes-v3-2-8e4adb1619e6@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
Changes in v4:
- Gate the dump in zap_threads() on freezing() and JOBCTL_TRAP_FREEZE and
  let dump_interrupted() test the flag, as suggested by Oleg.
- Drop the signal_pending() hook and coredump_signal_pending().
- Blame the cgroup v2 freezer commit as well.
- Link to v3: https://patch.msgid.link/20260921-work-coredump-fixes-v3-2-8e4adb1619e6@kernel.org

 fs/coredump.c | 13 +++++--------
 1 file changed, 5 insertions(+), 8 deletions(-)

diff --git a/fs/coredump.c b/fs/coredump.c
index d88a65e37ee1..af09a4660341 100644
--- a/fs/coredump.c
+++ b/fs/coredump.c
@@ -508,7 +508,9 @@ static int zap_threads(struct task_struct *tsk,
 	int nr = -EAGAIN;
 
 	spin_lock_irq(&tsk->sighand->siglock);
-	if (!(signal->flags & SIGNAL_GROUP_EXIT) && !signal->group_exec_task) {
+	/* A freeze requested before the dump would be lost with TIF_SIGPENDING. */
+	if (!(signal->flags & SIGNAL_GROUP_EXIT) && !signal->group_exec_task &&
+	    !freezing(tsk) && !(tsk->jobctl & JOBCTL_TRAP_FREEZE)) {
 		/* Allow SIGKILL, see prepare_signal() */
 		signal->core_state = core_state;
 		nr = zap_process(signal, exit_code);
@@ -580,13 +582,8 @@ static void coredump_finish(bool core_dumped)
 
 static bool dump_interrupted(void)
 {
-	/*
-	 * SIGKILL or freezing() interrupt the coredumping. Perhaps we
-	 * can do try_to_freeze() and check __fatal_signal_pending(),
-	 * but then we need to teach dump_write() to restart and clear
-	 * TIF_SIGPENDING.
-	 */
-	return fatal_signal_pending(current) || freezing(current);
+	/* Only SIGKILL and the freezers set it after zap_threads(). */
+	return task_sigpending(current);
 }
 
 static void wait_for_dump_helpers(struct file *file)
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* Re: [PATCH v3 08/17] exit: hang up the tty before closing the files
  2026-09-21 13:44 ` [PATCH v3 08/17] exit: hang up the tty before closing the files Christian Brauner
@ 2026-09-24 12:10   ` Oleg Nesterov
  0 siblings, 0 replies; 29+ messages in thread
From: Oleg Nesterov @ 2026-09-24 12:10 UTC (permalink / raw)
  To: Christian Brauner
  Cc: Chris Mason, linux-fsdevel, Jens Axboe, Alexander Viro, Jan Kara,
	NeilBrown, Ingo Molnar, Peter Zijlstra, linux-mm, io-uring

On 09/21, Christian Brauner wrote:
>
> --- a/kernel/exit.c
> +++ b/kernel/exit.c
> @@ -1003,10 +1003,11 @@ void __noreturn do_exit(long code)
>
>  	exit_sem(tsk);
>  	exit_shm(tsk);
> -	exit_files(tsk);
> -	exit_fs(tsk);
> +	/* Hang the tty up before the last close of it can clear the session. */
>  	if (group_dead)
>  		disassociate_ctty(1);
> +	exit_files(tsk);
> +	exit_fs(tsk);

Acked-by: Oleg Nesterov <oleg@redhat.com>


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v3 12/17] exec: cancel io_uring requests before de_thread()
  2026-09-21 13:45 ` [PATCH v3 12/17] exec: cancel io_uring requests before de_thread() Christian Brauner
  2026-09-21 14:14   ` Oleg Nesterov
@ 2026-09-24 14:19   ` Jens Axboe
  1 sibling, 0 replies; 29+ messages in thread
From: Jens Axboe @ 2026-09-24 14:19 UTC (permalink / raw)
  To: Christian Brauner, Oleg Nesterov, Chris Mason, linux-fsdevel
  Cc: Alexander Viro, Jan Kara, NeilBrown, Ingo Molnar, Peter Zijlstra,
	linux-mm, io-uring, stable

On 9/21/26 7:45 AM, Christian Brauner wrote:
> begin_new_exec() cancels the exec'ing thread's io_uring requests after
> de_thread() made it the only thread in the thread-group. Now
> io_uring_task_cancel() runs task work after. If it runs a
> create_worker_cb() item then a io-wq queued on the thread before the
> exec creates a new io-wq worker thread.
> 
> Said thread joins the thread-group which de_thread() just took down. The
> io-iwq sticks around until it notices that its workqueue is dying. But
> code after de_thread() relies on being single-threaded.
> 
> Move io_uring_task_cancel() before de_thread() but behind the point of
> no return. Any worker created during task work run from
> io_uring_task_cancel() is just a sibling thread that de_thread() will
> take down and reap.
> 
> Once io_uring_task_cancel() returned the thread has neither a task
> context nor a workqueue left so nothing can add a thread behind
> de_thread()'s back anymore.

Reviewed-by: Jens Axboe <axboe@kernel.dk>

-- 
Jens Axboe


^ permalink raw reply	[flat|nested] 29+ messages in thread

end of thread, other threads:[~2026-09-24 14:19 UTC | newest]

Thread overview: 29+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-21 13:44 [PATCH v3 00/17] coredump & signals: an impossible affair Christian Brauner
2026-09-21 13:44 ` [PATCH v3 01/17] coredump: hold RCU while releasing parked threads Christian Brauner
2026-09-21 13:44 ` [PATCH v3 02/17] signal: only SIGKILL and the freezers interrupt a coredumping task Christian Brauner
2026-09-22 12:55   ` Oleg Nesterov
2026-09-22 14:33     ` Christian Brauner
2026-09-21 13:44 ` [PATCH v3 03/17] coredump: parse a snapshot of core_pattern Christian Brauner
2026-09-21 13:44 ` [PATCH v3 04/17] io-wq: order the exit bit against worker creation task work Christian Brauner
2026-09-21 13:44 ` [PATCH v3 05/17] signal: don't retarget shared signals in a dying thread group Christian Brauner
2026-09-21 13:44 ` [PATCH v3 06/17] selftests/coredump: test shared signal retargeting during a dump Christian Brauner
2026-09-21 13:44 ` [PATCH v3 07/17] fork: release the files of a failed fork after sched_cancel_fork() Christian Brauner
2026-09-21 13:44 ` [PATCH v3 08/17] exit: hang up the tty before closing the files Christian Brauner
2026-09-24 12:10   ` Oleg Nesterov
2026-09-21 13:44 ` [PATCH v3 09/17] ptrace: refuse to change the signal mask of a user worker Christian Brauner
2026-09-21 14:14   ` Oleg Nesterov
2026-09-21 13:44 ` [PATCH v3 10/17] selftests/coredump: test a user worker as the coredumping thread Christian Brauner
2026-09-21 13:45 ` [PATCH v3 11/17] selftests/coredump: expect PTRACE_SETSIGMASK to be refused on a user worker Christian Brauner
2026-09-21 13:45 ` [PATCH v3 12/17] exec: cancel io_uring requests before de_thread() Christian Brauner
2026-09-21 14:14   ` Oleg Nesterov
2026-09-24 14:19   ` Jens Axboe
2026-09-21 13:45 ` [PATCH v3 13/17] fork: move the coredump and exec checks into create_io_thread() Christian Brauner
2026-09-21 14:15   ` Oleg Nesterov
2026-09-21 13:45 ` [PATCH v3 14/17] fork: don't create io threads once PF_POSTCOREDUMP is set Christian Brauner
2026-09-21 14:26   ` Oleg Nesterov
2026-09-21 13:45 ` [PATCH v3 15/17] fork: use SIG_KERNEL_ONLY_MASK for the user worker signal mask Christian Brauner
2026-09-21 14:29   ` Oleg Nesterov
2026-09-21 13:45 ` [PATCH v3 16/17] signal: enforce the user worker signal mask in __set_task_blocked() Christian Brauner
2026-09-21 16:16   ` Oleg Nesterov
2026-09-21 20:05     ` Christian Brauner
2026-09-21 13:45 ` [PATCH v3 17/17] fs: close files from the highest descriptor down Christian Brauner

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox