You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup从不同维基百科页面抓取剧集简介?

通用网页抓取适配方案

维基百科里不同剧的页面结构差异很大——有的单页就塞了所有剧集内容,有的得跳去剧集列表页,甚至还要再点进各季页面。要做通用脚本,得加几步结构检测和动态跳转逻辑:

  • 先检查当前页有没有剧集表
    先试着找wikiepisodetable类的表格,要是找不到,就定位页面里带「List of XXX episodes」字样的链接,自动跳过去。代码大概是这样:

    # 先找当前页面的剧集表
    tab = soup.find("table", {"class": "wikitable plainrowheaders wikiepisodetable"})
    if not tab:
        # 找剧集列表链接
        episode_list_link = soup.find("a", string=re.compile(r"List of.*episodes"))
        if episode_list_link:
            # 拼完整URL跳转
            list_url = "https://en.wikipedia.org" + episode_list_link["href"]
            source = urlopen(list_url).read()
            soup = BeautifulSoup(source, 'lxml')
            tab = soup.find("table", {"class": "wikitable plainrowheaders wikiepisodetable"})
    
  • 处理分季页面的情况
    要是跳转后的页面是分季汇总(比如《生活大爆炸》的列表页),就得遍历每个季的链接,挨个抓内容:

    # 找所有季的链接
    season_links = soup.find_all("a", string=re.compile(r"Season \d+"))
    if season_links:
        all_episode_text = ""
        for link in season_links:
            season_url = "https://en.wikipedia.org" + link["href"]
            season_source = urlopen(season_url).read()
            season_soup = BeautifulSoup(season_source, 'lxml')
            season_tab = season_soup.find("table", {"class": "wikitable plainrowheaders wikiepisodetable"})
            # 提取当前季的简介(复用你原来的逻辑)
            spans = season_tab.find_all('td')
            x = [i for i in range(4, len(spans), 5)]
            tds = [spans[i] for i in x]
            season_text = ''.join([td.text for td in tds])
            # 清理文本
            season_text = re.sub(r'\[.*?\]+', '', season_text).replace('\n', '')
            all_episode_text += season_text + "\n\n"
        text = all_episode_text
    else:
        # 单页有完整剧集表,直接用原逻辑提取
        spans = tab.find_all('td')
        x = [i for i in range(4, len(spans), 5)]
        tds = [spans[i] for i in x]
        text = ''.join([td.text for td in tds])
        text = re.sub(r'\[.*?\]+', '', text).replace('\n', '')
    
  • 再加个鲁棒性优化
    别用固定索引找简介列,换成通过表头文本定位——比如找表头里的「Plot」「Summary」「Synopsis」,这样就算表格列顺序变了也能拿到正确内容:

    # 动态找简介列的位置
    headers = tab.find_all("th")
    plot_col_index = None
    for idx, header in enumerate(headers):
        if header.text.strip() in ["Plot", "Summary", "Synopsis"]:
            plot_col_index = idx
            break
    # 提取对应列的内容
    if plot_col_index is not None:
        tds = [row.find_all('td')[plot_col_index] for row in tab.find_all("tr")[1:]]  # 跳过表头行
    
更省心的方案:用维基百科官方API/数据库

网页抓取太容易受页面结构变动影响,不如直接用官方提供的结构化数据渠道:

  • MediaWiki API
    直接通过API查剧集的内容,不用解析HTML。比如查《欧比旺》的剧集数据:

    import requests
    
    api_url = "https://en.wikipedia.org/w/api.php"
    params = {
        "action": "query",
        "format": "json",
        "titles": "Obi-Wan Kenobi (TV series)",
        "prop": "revisions",
        "rvprop": "content",
        "rvparse": "true"  # 返回解析好的HTML,方便提取
    }
    response = requests.get(api_url, params=params)
    data = response.json()
    # 从返回的HTML里提取剧集表就行,数据比网页稳定多了
    

    进阶点还能查Wikidata的实体数据,直接拿到结构化的剧集关联关系。

  • 维基数据(Wikidata)
    这是维基的结构化知识库,每个剧和剧集都有唯一的QID。比如《安多》的QID是Q106443554,你可以:

    1. 通过剧名查到对应的QID;
    2. 查这个QID关联的所有剧集条目(属性P1542);
    3. 挨个获取每个剧集的简介内容(属性P1366)或者维基页面摘要。
  • 预导出数据集
    要是你需要批量处理大量剧集,可以下载维基百科的XML dump,用WikiExtractor这类工具提取结构化内容,本地处理效率更高,也不用怕网页结构变了。


内容的提问来源于stack exchange,提问作者Salvo R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 22:15:31