如何使用asyncio将我的Python爬虫程序改为异步执行?
异步并发优化方案:多链接同时爬取
没问题!要把你的同步爬取改成异步多链接并发爬取,咱们可以用asyncio搭配aiohttp(专门适配异步场景的HTTP库,比同步的requests更适合)来实现,这样就能同时发起多个请求,不用等前一个任务完成再处理下一个,效率会提升很多。
第一步:先安装依赖
因为要用到异步HTTP客户端,先装一下需要的包:
pip install aiohttp beautifulsoup4 lxml
修改后的完整代码
import urllib import re import asyncio import aiohttp from bs4 import BeautifulSoup address = 'https://google.com/search?q=' # Default Google search address start file = open( "OCR.txt", "rt" ) # Open text document that contains the question word = file.read() file.close() myList = [item for item in word.split('\n')] newString = ' '.join(myList) # The question is on multiple lines so this joins them together with proper spacing qstr = urllib.parse.quote_plus(newString) # Encode the string newWord = address + qstr # Combine the base and the encoded query # 读取答案选项 answers = open("ocr2.txt", "rt") ansTable = answers.read() answers.close() ans = ansTable.splitlines() ans1 = str(ans[0]) ans2 = str(ans[2]) ans3 = str(ans[4]) ans1Score = 0 ans2Score = 0 ans3Score = 0 links = [] # 异步函数:爬取单个链接并统计答案出现次数 async def fetch_and_count(session, url): global ans1Score, ans2Score, ans3Score try: if url.endswith('pdf'): return async with session.get(url) as response: text = await response.text() soup2 = BeautifulSoup(text, 'lxml') for p in soup2.find_all('p'): extraBlock = str(p) ans1Score += extraBlock.count(ans1) ans2Score += extraBlock.count(ans2) ans3Score += extraBlock.count(ans3) except Exception as e: print(f"爬取链接 {url} 时出错: {e}") async def main(): global ans1Score, ans2Score, ans3Score, links # 先获取Google搜索结果 async with aiohttp.ClientSession() as session: response = await session.get(newWord) soup = BeautifulSoup(await response.text(), 'lxml') # 解析搜索结果中的链接 for r in soup.find_all(class_='r'): linkRaw = str(r) match = re.search("(?P<url>https?://[^\s]+)", linkRaw) if match: link = match.group("url") if '&' in link: link = link.split('&')[0] links.append(link) # 遍历搜索结果区块,统计初始分数 ans1Found = False ans2Found = False ans3Found = False for g in soup.find_all(class_='g'): webBlock = str(g) ans1Tally = webBlock.count(ans1) ans2Tally = webBlock.count(ans2) ans3Tally = webBlock.count(ans3) if ans1Tally > 0: ans1Score += ans1Tally ans1Found = True if ans2Tally > 0: ans2Score += ans2Tally ans2Found = True if ans3Tally > 0: ans3Score += ans3Tally ans3Found = True # 如果有答案没在搜索摘要里找到,就异步爬取链接内容 if not (ans1Found and ans2Found and ans3Found): # 创建异步会话,批量发起请求 async with aiohttp.ClientSession() as session: # 把要爬的链接转成异步任务列表 tasks = [fetch_and_count(session, link) for link in links] # 并发执行所有任务 await asyncio.gather(*tasks) # 保存结果到文件 with open("Results.txt", "w") as results: results.write(newString + '\n\n') results.write(f"{ans1}: {ans1Score}\n") results.write(f"{ans2}: {ans2Score}\n") results.write(f"{ans3}: {ans3Score}") # 打印结果 print(' ') print('-----') print(f"{ans1}: {ans1Score}") print(f"{ans2}: {ans2Score}") print(f"{ans3}: {ans3Score}") print('-----') if __name__ == "__main__": asyncio.run(main())
关键改动说明
- 替换同步HTTP库:把原来的
requests换成aiohttp,用异步的ClientSession来发起请求,避免阻塞 - 封装异步任务:把单个链接的爬取和统计逻辑封装成
fetch_and_count异步函数,每个链接对应一个独立任务 - 并发执行任务:用
asyncio.gather(*tasks)一次性启动所有爬取任务,让它们同时运行,不用等待前一个完成 - 全局分数变量:因为异步任务需要共享分数统计,这里用了全局变量(如果想更优雅可以用类或者字典来传递,不过小脚本用全局变量足够简单)
- 异常处理:给异步爬加了异常捕获,避免单个链接爬取失败导致整个程序崩溃
这样修改后,你的程序就能同时爬取多个链接,大幅提升效率啦!
内容的提问来源于stack exchange,提问作者DevinGP
相关产品推荐
相关产品推荐

