You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用asyncio将我的Python爬虫程序改为异步执行?

异步并发优化方案:多链接同时爬取

没问题!要把你的同步爬取改成异步多链接并发爬取,咱们可以用asyncio搭配aiohttp(专门适配异步场景的HTTP库,比同步的requests更适合)来实现,这样就能同时发起多个请求,不用等前一个任务完成再处理下一个,效率会提升很多。

第一步:先安装依赖

因为要用到异步HTTP客户端,先装一下需要的包:

pip install aiohttp beautifulsoup4 lxml

修改后的完整代码

import urllib
import re
import asyncio
import aiohttp
from bs4 import BeautifulSoup

address = 'https://google.com/search?q=' # Default Google search address start
file = open( "OCR.txt", "rt" ) # Open text document that contains the question
word = file.read()
file.close()
myList = [item for item in word.split('\n')]
newString = ' '.join(myList) # The question is on multiple lines so this joins them together with proper spacing
qstr = urllib.parse.quote_plus(newString) # Encode the string
newWord = address + qstr # Combine the base and the encoded query

# 读取答案选项
answers = open("ocr2.txt", "rt")
ansTable = answers.read()
answers.close()
ans = ansTable.splitlines()
ans1 = str(ans[0])
ans2 = str(ans[2])
ans3 = str(ans[4])
ans1Score = 0
ans2Score = 0
ans3Score = 0
links = []

# 异步函数:爬取单个链接并统计答案出现次数
async def fetch_and_count(session, url):
    global ans1Score, ans2Score, ans3Score
    try:
        if url.endswith('pdf'):
            return
        async with session.get(url) as response:
            text = await response.text()
            soup2 = BeautifulSoup(text, 'lxml')
            for p in soup2.find_all('p'):
                extraBlock = str(p)
                ans1Score += extraBlock.count(ans1)
                ans2Score += extraBlock.count(ans2)
                ans3Score += extraBlock.count(ans3)
    except Exception as e:
        print(f"爬取链接 {url} 时出错: {e}")

async def main():
    global ans1Score, ans2Score, ans3Score, links
    # 先获取Google搜索结果
    async with aiohttp.ClientSession() as session:
        response = await session.get(newWord)
        soup = BeautifulSoup(await response.text(), 'lxml')
    
    # 解析搜索结果中的链接
    for r in soup.find_all(class_='r'):
        linkRaw = str(r)
        match = re.search("(?P<url>https?://[^\s]+)", linkRaw)
        if match:
            link = match.group("url")
            if '&amp;' in link:
                link = link.split('&amp;')[0]
            links.append(link)
    
    # 遍历搜索结果区块,统计初始分数
    ans1Found = False
    ans2Found = False
    ans3Found = False
    for g in soup.find_all(class_='g'):
        webBlock = str(g)
        ans1Tally = webBlock.count(ans1)
        ans2Tally = webBlock.count(ans2)
        ans3Tally = webBlock.count(ans3)
        if ans1Tally > 0:
            ans1Score += ans1Tally
            ans1Found = True
        if ans2Tally > 0:
            ans2Score += ans2Tally
            ans2Found = True
        if ans3Tally > 0:
            ans3Score += ans3Tally
            ans3Found = True
    
    # 如果有答案没在搜索摘要里找到,就异步爬取链接内容
    if not (ans1Found and ans2Found and ans3Found):
        # 创建异步会话,批量发起请求
        async with aiohttp.ClientSession() as session:
            # 把要爬的链接转成异步任务列表
            tasks = [fetch_and_count(session, link) for link in links]
            # 并发执行所有任务
            await asyncio.gather(*tasks)
    
    # 保存结果到文件
    with open("Results.txt", "w") as results:
        results.write(newString + '\n\n')
        results.write(f"{ans1}: {ans1Score}\n")
        results.write(f"{ans2}: {ans2Score}\n")
        results.write(f"{ans3}: {ans3Score}")
    
    # 打印结果
    print(' ')
    print('-----')
    print(f"{ans1}: {ans1Score}")
    print(f"{ans2}: {ans2Score}")
    print(f"{ans3}: {ans3Score}")
    print('-----')

if __name__ == "__main__":
    asyncio.run(main())

关键改动说明

  • 替换同步HTTP库:把原来的requests换成aiohttp,用异步的ClientSession来发起请求,避免阻塞
  • 封装异步任务:把单个链接的爬取和统计逻辑封装成fetch_and_count异步函数,每个链接对应一个独立任务
  • 并发执行任务:用asyncio.gather(*tasks)一次性启动所有爬取任务,让它们同时运行,不用等待前一个完成
  • 全局分数变量:因为异步任务需要共享分数统计,这里用了全局变量(如果想更优雅可以用类或者字典来传递,不过小脚本用全局变量足够简单)
  • 异常处理:给异步爬加了异常捕获,避免单个链接爬取失败导致整个程序崩溃

这样修改后,你的程序就能同时爬取多个链接,大幅提升效率啦!

内容的提问来源于stack exchange,提问作者DevinGP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:41:33