You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python Wikipedia包并行获取3个维基百科页面正文以提升性能?

并行获取维基百科页面内容的实用方案

当然有办法解决这个性能问题!这类网络IO密集型任务,并行请求绝对是提升效率的关键——毕竟你大部分时间都在等服务器响应,而不是占用CPU计算。下面给你两种Python生态里常用的实现方式,直接就能用:

方案一:用线程池快速改造现有代码

如果你想尽量复用现有的wikipedia包代码,concurrent.futures.ThreadPoolExecutor是最省心的选择。线程池可以同时启动多个线程去请求维基百科,不用等前一个请求完成再启动下一个。

示例代码:

import wikipedia
from concurrent.futures import ThreadPoolExecutor

# 封装获取单篇内容的函数,包含异常处理
def fetch_wiki_page(title):
    try:
        page = wikipedia.page(title)
        # 返回标题和内容摘要(你可以根据需求调整返回格式)
        return f"【{page.title}】\n{page.content[:600]}..."
    except wikipedia.exceptions.PageError:
        return f"⚠️ 页面「{title}」不存在"
    except Exception as e:
        return f"❌ 获取「{title}」失败: {str(e)}"

# 要获取的3个页面标题
target_titles = ["Python (programming language)", "Artificial intelligence", "Machine learning"]

if __name__ == "__main__":
    # 初始化线程池,max_workers设为3刚好匹配你的任务数
    with ThreadPoolExecutor(max_workers=3) as executor:
        # 批量提交任务并获取结果
        results = list(executor.map(fetch_wiki_page, target_titles))
    
    # 打印结果
    for res in results:
        print("\n" + "-"*60 + "\n")
        print(res)

为什么这个方案好用?

  • 几乎不用改你原来的逻辑,只是把串行调用改成了线程池批量提交
  • 线程池会自动管理线程的创建和销毁,不用手动处理复杂的多线程逻辑
  • 对于3个任务来说,性能提升非常明显——原来要等3次网络耗时,现在差不多只需要1次的时间

方案二:用异步IO实现更高效的请求

如果以后你要处理更多页面,异步IO会是更优的选择。不过wikipedia包本身是基于同步的requests库,所以我们可以用aiohttp发起异步请求,再用BeautifulSoup解析页面内容。

示例代码:

import asyncio
import aiohttp
from bs4 import BeautifulSoup

async def fetch_wiki_async(session, title):
    # 构造维基百科页面的URL
    page_url = f"https://en.wikipedia.org/wiki/{title.replace(' ', '_')}"
    try:
        async with session.get(page_url) as resp:
            if resp.status != 200:
                return f"⚠️ 请求「{title}」失败,状态码: {resp.status}"
            # 异步获取页面HTML
            html_content = await resp.text()
            # 解析正文内容
            soup = BeautifulSoup(html_content, "html.parser")
            content_div = soup.find(id="mw-content-text")
            # 提取所有段落文本
            paragraphs = [p.get_text().strip() for p in content_div.find_all("p") if p.get_text().strip()]
            full_content = "\n".join(paragraphs)
            return f"【{title}】\n{full_content[:600]}..."
    except Exception as e:
        return f"❌ 获取「{title}」失败: {str(e)}"

async def main():
    target_titles = ["Python (programming language)", "Artificial intelligence", "Machine learning"]
    # 创建异步会话
    async with aiohttp.ClientSession() as session:
        # 创建所有异步任务
        tasks = [fetch_wiki_async(session, title) for title in target_titles]
        # 等待所有任务完成并获取结果
        results = await asyncio.gather(*tasks)
    
    # 输出结果
    for res in results:
        print("\n" + "-"*60 + "\n")
        print(res)

if __name__ == "__main__":
    asyncio.run(main())

注意事项

  • 这个方案需要先安装依赖:pip install aiohttp beautifulsoup4
  • 异步IO在单线程内处理多个请求,资源占用比线程池更低,适合大规模任务
  • 不管用哪种方案,都不要设置过高的并发数——维基百科有反爬机制,过于频繁的请求可能会被暂时封禁IP

内容的提问来源于stack exchange,提问作者delhics

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:49:36