You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在同时运行两个网页爬虫时优先获取首个结果并用于后续处理

解决方案

因为爬虫属于IO密集型任务(大部分时间在等待网络响应),用threading确实比多进程更高效,线程开销更小。下面给你两种简洁的实现方式:

方式一:用threading+回调逻辑

直接让每个线程跑完爬虫后立刻处理结果,哪个先完成就先发推特:

import threading
from scraper1 import main_1
from scraper2 import main_2
from twitter import post_tweet

def run_and_post(scraper_func, *args):
    data = scraper_func(*args)
    post_tweet(f"New data is {data}")

if __name__ == '__main__':
    # 启动两个线程分别执行爬虫
    t1 = threading.Thread(target=run_and_post, args=(main_1, 'www.website1.com', 'June'))
    t2 = threading.Thread(target=run_and_post, args=(main_2,))
    
    t1.start()
    t2.start()
    
    # 可选:让主线程等两个爬虫都结束再退出
    t1.join()
    t2.join()

这个逻辑非常直白:每个线程独立跑爬虫,拿到结果后直接调用发推特的函数,完全不用等待另一个线程完成。

方式二:优化你的multiprocessing代码

如果想继续用多进程,也可以通过apply_async的callback参数实现“拿到结果就处理”,不用手动调用阻塞的get():

from multiprocessing import Pool
from scraper1 import main_1
from scraper2 import main_2
from twitter import post_tweet

def handle_result(data):
    post_tweet(f"New data is {data}")

if __name__ == '__main__':
    with Pool(processes=2) as pool:
        # 给每个异步任务绑定回调,结果返回后自动执行发推特逻辑
        pool.apply_async(main_1, args=('www.website1.com','June'), callback=handle_result)
        pool.apply_async(main_2, args=(), callback=handle_result)
        
        # 等待所有任务完成,避免Pool提前关闭
        pool.close()
        pool.join()

这里的callback参数会在任务结束后自动触发,传入爬虫返回的数据,两个爬虫谁先跑完谁就先发推特,不用等另一个。

为什么原来的代码不行?

你之前的r1.get()会强制阻塞主线程,必须等第一个爬虫跑完才会去拿第二个的结果,所以只能等两个都完成后才能处理。上面两种方案都是让任务完成后自动触发处理逻辑,彻底避免了阻塞等待。

内容的提问来源于stack exchange,提问作者appletree3

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 14:36:35