如何理解FutureTask.removeWaiter()中‘check for race’注释及重启竞争的选择?
解析JDK 1.7 FutureTask.removeWaiter()中的竞争处理逻辑
原方法代码及注释
/** * Tries to unlink a timed-out or interrupted wait node to avoid * accumulating garbage. Internal nodes are simply unspliced * without CAS since it is harmless if they are traversed anyway * by releasers. To avoid effects of unsplicing from already * removed nodes, the list is retraversed in case of an apparent * race. This is slow when there are a lot of nodes, but we don't * expect lists to be long enough to outweigh higher-overhead * schemes. */ private void removeWaiter(WaitNode node) { if (node != null) { node.thread = null; retry: for (;;) { // restart on removeWaiter race for (WaitNode pred = null, q = waiters, s; q != null; q = s) { s = q.next; if (q.thread != null) pred = q; else if (pred != null) { pred.next = s; if (pred.thread == null) // check for race continue retry; } else if (!UNSAFE.compareAndSwapObject(this, waitersOffset, q, s)) continue retry; } break; } } }
用户的理解与疑问
- 竞争条件:当线程成功找到pred节点时,若其他线程尝试移除pred节点,就会发生竞争。
- 个人观点:
q.next、q.thread和pred.thread均为volatile,因此当前线程即使继续与其他线程竞争,结果依然有效。 - 疑问:为何Doug Lea选择重启竞争而非继续竞争?或是我的理解存在错误?
为什么选择重启竞争而非继续?
咱们先拆解清楚这里的核心逻辑:
首先你对volatile变量的可见性判断是完全正确的——这些字段的volatile修饰确实保证了线程间的状态可见性,但问题的核心不在可见性,而在链表结构的一致性。
当代码走到if (pred.thread == null)时,说明我们刚刚用来拼接的pred节点本身已经被标记为待移除状态(thread置为null)了。这时候我们之前做的pred.next = s操作已经失去意义:因为pred很快会被其他线程从链表中摘掉,这个修改不会被后续的遍历或操作感知到,甚至可能导致链表出现断链、残留无效节点的情况。
如果选择继续当前的遍历,我们可能会跳过后续需要处理的节点,毕竟此时的pred已经是无效节点,基于它的遍历路径是不可靠的。而重启遍历(continue retry)相当于从头开始重新梳理整个等待链表,确保我们能准确找到所有需要移除的节点,保证链表结构的完整性和正确性。
Doug Lea之所以这么设计,还有一个现实考量:FutureTask的等待链表通常不会很长(毕竟大部分任务会很快完成,不会有大量线程长期阻塞等待),所以偶尔的重启遍历带来的性能开销完全可以接受。相比之下,这种简单粗暴的重启逻辑比复杂的“继续竞争”逻辑更不容易出错,维护成本也更低——这完全符合他在并发工具类设计中“简单可靠优先,兼顾性能”的风格。
内容的提问来源于stack exchange,提问作者Hello World
相关产品推荐
相关产品推荐

