如何使用Python Wikipedia包并行获取3个维基百科页面正文以提升性能?
并行获取维基百科页面内容的实用方案
当然有办法解决这个性能问题!这类网络IO密集型任务,并行请求绝对是提升效率的关键——毕竟你大部分时间都在等服务器响应,而不是占用CPU计算。下面给你两种Python生态里常用的实现方式,直接就能用:
方案一:用线程池快速改造现有代码
如果你想尽量复用现有的wikipedia包代码,concurrent.futures.ThreadPoolExecutor是最省心的选择。线程池可以同时启动多个线程去请求维基百科,不用等前一个请求完成再启动下一个。
示例代码:
import wikipedia from concurrent.futures import ThreadPoolExecutor # 封装获取单篇内容的函数,包含异常处理 def fetch_wiki_page(title): try: page = wikipedia.page(title) # 返回标题和内容摘要(你可以根据需求调整返回格式) return f"【{page.title}】\n{page.content[:600]}..." except wikipedia.exceptions.PageError: return f"⚠️ 页面「{title}」不存在" except Exception as e: return f"❌ 获取「{title}」失败: {str(e)}" # 要获取的3个页面标题 target_titles = ["Python (programming language)", "Artificial intelligence", "Machine learning"] if __name__ == "__main__": # 初始化线程池,max_workers设为3刚好匹配你的任务数 with ThreadPoolExecutor(max_workers=3) as executor: # 批量提交任务并获取结果 results = list(executor.map(fetch_wiki_page, target_titles)) # 打印结果 for res in results: print("\n" + "-"*60 + "\n") print(res)
为什么这个方案好用?
- 几乎不用改你原来的逻辑,只是把串行调用改成了线程池批量提交
- 线程池会自动管理线程的创建和销毁,不用手动处理复杂的多线程逻辑
- 对于3个任务来说,性能提升非常明显——原来要等3次网络耗时,现在差不多只需要1次的时间
方案二:用异步IO实现更高效的请求
如果以后你要处理更多页面,异步IO会是更优的选择。不过wikipedia包本身是基于同步的requests库,所以我们可以用aiohttp发起异步请求,再用BeautifulSoup解析页面内容。
示例代码:
import asyncio import aiohttp from bs4 import BeautifulSoup async def fetch_wiki_async(session, title): # 构造维基百科页面的URL page_url = f"https://en.wikipedia.org/wiki/{title.replace(' ', '_')}" try: async with session.get(page_url) as resp: if resp.status != 200: return f"⚠️ 请求「{title}」失败,状态码: {resp.status}" # 异步获取页面HTML html_content = await resp.text() # 解析正文内容 soup = BeautifulSoup(html_content, "html.parser") content_div = soup.find(id="mw-content-text") # 提取所有段落文本 paragraphs = [p.get_text().strip() for p in content_div.find_all("p") if p.get_text().strip()] full_content = "\n".join(paragraphs) return f"【{title}】\n{full_content[:600]}..." except Exception as e: return f"❌ 获取「{title}」失败: {str(e)}" async def main(): target_titles = ["Python (programming language)", "Artificial intelligence", "Machine learning"] # 创建异步会话 async with aiohttp.ClientSession() as session: # 创建所有异步任务 tasks = [fetch_wiki_async(session, title) for title in target_titles] # 等待所有任务完成并获取结果 results = await asyncio.gather(*tasks) # 输出结果 for res in results: print("\n" + "-"*60 + "\n") print(res) if __name__ == "__main__": asyncio.run(main())
注意事项
- 这个方案需要先安装依赖:
pip install aiohttp beautifulsoup4 - 异步IO在单线程内处理多个请求,资源占用比线程池更低,适合大规模任务
- 不管用哪种方案,都不要设置过高的并发数——维基百科有反爬机制,过于频繁的请求可能会被暂时封禁IP
内容的提问来源于stack exchange,提问作者delhics
相关产品推荐
相关产品推荐

