You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python multiprocessing.Process调用网络后无法终止问题求助

多进程爬取进程无法正常退出的解决方案

问题根源分析

你的问题主要来自三个方面:

  1. 共享状态管理不当:用Queue维护COUNT和paperNumber这类共享计数,频繁的get/put操作会引发进程间竞争,出现计数错误或死锁,最终导致进程无法退出。
  2. 网络连接资源未彻底释放:手动调用response.close()无法确保连接被可靠回收,存在资源泄漏风险。
  3. 进程退出逻辑有缺陷:多个Read进程同时操作COUNT,会出现重复计数或无法触发退出条件的情况,导致进程陷入无限循环。

具体修复方案

1. 替换共享状态实现

将用于计数的Queue替换为multiprocessing.Manager提供的Value,它专门用于多进程间的数值共享,避免Queue的锁竞争问题。

2. 用上下文管理器管理requests连接

使用with requests.get(...) as response的方式,确保请求完成后连接自动释放,无需手动调用close()。

3. 优化进程退出逻辑

  • 写入进程完成任务后,向队列放入None作为结束标记,每个读取进程对应一个标记
  • 读取进程检测到结束标记时主动退出,不再依赖有问题的COUNT计数逻辑

4. 避免跨进程通信死锁

使用Manager.Queue替代普通Queue,它更适合跨进程场景,降低死锁概率。

修改后的完整代码

from multiprocessing import Process, Manager
import time
import requests
import re
from bs4 import BeautifulSoup

headers = {
    'user-agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.129 Safari/537.36,Mozilla/5.0 (Macintosh; Intel Mac OS X 10_7_5) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/27.0.1453.93 Safari/537.36",
    'Connection': 'close'
}

## 改用上下文管理器获取网页内容
def GetUrlInfo(url):
    with requests.get(url=url, headers=headers) as response:
        response.encoding = 'utf-8'
        SoupData = BeautifulSoup(response.text, 'lxml')
        return SoupData

def GetVolumeUrlfromUrl(url:str)->str:
    """input is Journal's url and output is a link and a text description to each issue of the journal"""
    url = re.sub('http:', 'https:', url)
    SoupDataTemp = GetUrlInfo(url+'index.html')
    SoupData = SoupDataTemp.find_all('li')
    UrlALL = []
    for i in SoupData:
        if i.find('a') != None:
            volumeUrlRule = '<a href=\"(.*?)\">(.*?)</a>'
            volumeUrlTemp = re.findall(volumeUrlRule,str(i),re.I)
            for u in volumeUrlTemp:
                if re.findall(url, u[0]):
                    UrlALL.append((u[0], u[1]), )
    return UrlALL

def GetPaperBaseInfoFromUrlAll(url:str)->str:
    """The input is the url and the output is all the paper information obtained from the web page,
    including, doi, title, author, and the date about this volume """
    soup = GetUrlInfo(url)
    temp1 = soup.find_all('li',class_='entry article')
    temp2= soup.find_all('h2')
    temp2=re.sub('\\n',' ',temp2[1].text)
    volumeYear = re.split(' ',temp2)[-1]
    paper = []
    for i in temp1:
        if i.find('div',class_='head').find('a')== None:
            paperDoi = ''
        else:
            paperDoi = i.find('div',class_='head').find('a')['href']
        title = i.find('cite').find('span',class_='title').text[:-2]
        paper.append([paperDoi,title])
    return paper,volumeYear

# 写入任务队列,添加结束标记
def Write(query, value, process_num):
    for item in value:
        query.put(item[0])
    # 给每个Read进程放一个结束标记
    for _ in range(process_num):
        query.put(None)
    print('write end')

# 读取任务队列并处理
def Read(query, paper_info, paper_count, process_id):
    while True:
        url = query.get()
        if url is None:
            # 收到结束标记,退出进程
            print(f"进程 {process_id} 结束")
            break
        try:
            paper, this_year = GetPaperBaseInfoFromUrlAll(url)
            print(f"connected {process_id} : {url}")
            # 更新共享计数,加锁保证原子性
            with paper_count.get_lock():
                paper_count.value += len(paper)
            paper_info.put((paper, this_year))
            print(f"进程 {process_id} 完成处理: {url}")
        except Exception as e:
            print(f"进程 {process_id} 处理失败 {url}: {str(e)}")
    print(f'read end {process_id}')

# 打印结果
def GetPaperInfo(paper_info, paper_count):
    print(f"总共爬取到 {paper_count.value} 篇论文")
    while not paper_info.empty():
        value = paper_info.get()
        print(value)

if __name__=='__main__':
    r_num = 10  # 读取进程数
    w_num = 1  # 写入进程数
    url = 'http://dblp.uni-trier.de/db/journals/talg/'
    UrlALL = GetVolumeUrlfromUrl(url)
    UrlLen = len(UrlALL)

    # 使用Manager创建跨进程共享的对象
    manager = Manager()
    q = manager.Queue()
    paper_info = manager.Queue()
    paper_count = manager.Value('i', 0)  # 'i'表示整数类型

    r_list = [Process(target=Read, args=(q, paper_info, paper_count, i)) for i in range(r_num)]
    w_list = [Process(target=Write, args=(q, UrlALL, r_num))]

    time_start = time.time()
    [task.start() for task in w_list]
    [task.start() for task in r_list]

    [task.join() for task in w_list]
    [task.join() for task in r_list]

    time_used = time.time() - time_start
    GetPaperInfo(paper_info, paper_count)
    print(f'time_used:{time_used}s')

关键修改说明

  • 用Manager.Value替代Queue管理计数,确保计数准确且无锁竞争
  • 用with语句管理requests连接,彻底释放网络资源
  • 写入进程完成后向队列放入None作为结束标记,每个读取进程收到后直接退出
  • 使用Manager.Queue替代普通Queue,避免跨进程通信死锁
  • 添加异常捕获,避免单个请求失败导致进程崩溃

内容的提问来源于stack exchange,提问作者Jack August

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 17:15:34