You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup抓取维基百科特定板块内的多元素

问题描述

我是从C#转Python的新手,想要抓取维基百科微软页面中历史板块的全部文本内容。已知Beautiful Soup不支持XPath,该板块的第一个元素为:

<div role="note" class="hatnote navigation-not-searchable">Main article: <a href="/wiki/History_of_Microsoft" title="History of Microsoft">History of Microsoft</a></div>

最后一个元素为:

<p>On January 18, 2022, Microsoft announced the acquisition of American video game developer and <a href="/wiki/Holding_company" title="Holding company">holding company</a> <a href="/wiki/Activision_Blizzard" title="Activision Blizzard">Activision Blizzard</a> in an all-cash deal worth $68.7 billion.<sup id="cite_ref-:0_150-0" class="reference"><a href="#cite_note-:0-150">[150]</a></sup> Activision Blizzard is best known for producing franchises, including but not limited to <i><a href="/wiki/Warcraft" title="Warcraft">Warcraft</a></i>, <i><a href="/wiki/Diablo_(series)" title="Diablo (series)">Diablo</a></i>, <i><a href="/wiki/Call_of_Duty" title="Call of Duty">Call of Duty</a></i>, <i><a href="/wiki/StarCraft" title="StarCraft">StarCraft</a></i>, <i><a href="/wiki/Candy_Crush_Saga" title="Candy Crush Saga">Candy Crush Saga</a></i>, <i><a href="/wiki/Crash_Bandicoot" title="Crash Bandicoot">Crash Bandicoot</a></i>, <i><a href="/wiki/Spyro" title="Spyro">Spyro the Dragon</a></i>, <i><a href="/wiki/Skylanders" title="Skylanders">Skylanders</a></i>, and <i><a href="/wiki/Overwatch_(video_game)" title="Overwatch (video game)">Overwatch</a></i>.<sup id="cite_ref-151" class="reference"><a href="#cite_note-151">[151]</a></sup> Activision and Microsoft each released statements saying the acquisition was to benefit their businesses in the <a href="/wiki/Metaverse" title="Metaverse">metaverse</a>, many saw Microsoft's acquisition of video game studios as an attempt to compete against <a href="/wiki/Meta_Platforms" title="Meta Platforms">Meta Platforms</a>, with <a href="/wiki/TheStreet" title="TheStreet">TheStreet</a> referring to Microsoft wanting to become "the <a href="/wiki/The_Walt_Disney_Company" title="The Walt Disney Company">Disney</a> of the metaverse".<sup id="cite_ref-152" class="reference"><a href="#cite_note-152">[152]</a></sup><sup id="cite_ref-153" class="reference"><a href="#cite_note-153">[153]</a></sup> Microsoft has not released statements regarding Activision's recent legal controversies regarding employee abuse, but reports have alleged that Activision CEO <a href="/wiki/Bobby_Kotick" title="Bobby Kotick">Bobby Kotick</a>, a major target of the controversy, will leave the company after the acquisition is finalized.<sup id="cite_ref-154" class="reference"><a href="#cite_note-154">[154]</a></sup> The deal is expected to close in 2023 followed by a review from the <a href="/wiki/US_Federal_Trade_Commission" class="mw-redirect" title="US Federal Trade Commission">US Federal Trade Commission</a>.<sup id="cite_ref-155" class="reference"><a href="#cite_note-155">[155]</a></sup><sup id="cite_ref-:0_150-1" class="reference"><a href="#cite_note-:0-150">[150]</a></sup></p>

请问如何获取这两个元素之间的所有信息?我已编写初始代码如下:

from bs4 import BeautifulSoup
import requests

url = 'https://en.wikipedia.org/wiki/Microsoft'
  
# call get method to request that page
page = requests.get(url)

soup = BeautifulSoup(page.text, "html.parser")
解决方案

你可以通过两种方式实现需求,一种是精准定位起始/结束元素并遍历中间节点,另一种是直接定位整个历史章节(更稳定)。

方法一:基于起始/结束元素抓取

通过属性和文本特征定位目标元素,再遍历两者之间的所有兄弟节点:

from bs4 import BeautifulSoup
import requests

url = 'https://en.wikipedia.org/wiki/Microsoft'
  
page = requests.get(url)
soup = BeautifulSoup(page.text, "html.parser")

# 定位起始元素:通过class和role组合匹配
start_element = soup.find('div', class_='hatnote navigation-not-searchable', role='note')

# 定位结束元素:通过段落中的标志性文本匹配
end_element = soup.find('p', string=lambda text: text and 'On January 18, 2022, Microsoft announced the acquisition' in text)

# 收集中间所有有效文本
history_content = []
current_element = start_element.next_sibling

while current_element != end_element:
    # 过滤空节点和无文本的元素
    if current_element and hasattr(current_element, 'get_text'):
        text = current_element.get_text(strip=True)
        if text:
            history_content.append(text)
    current_element = current_element.next_sibling

# 添加结束元素的文本
if end_element:
    history_content.append(end_element.get_text(strip=True))

# 输出结果
for content in history_content:
    print(content)

方法二:直接定位历史章节(推荐)

维基百科的章节结构固定,通过标题History定位整个板块,避免单个元素变动导致失效:

from bs4 import BeautifulSoup
import requests

url = 'https://en.wikipedia.org/wiki/Microsoft'
  
page = requests.get(url)
soup = BeautifulSoup(page.text, "html.parser")

history_section = []
# 遍历所有h2标题,找到History章节
for h2 in soup.find_all('h2'):
    if h2.get_text(strip=True) == 'History':
        # 收集h2之后的所有兄弟元素,直到下一个h2章节
        next_sib = h2.next_sibling
        while next_sib and next_sib.name != 'h2':
            if next_sib and hasattr(next_sib, 'get_text'):
                text = next_sib.get_text(strip=True)
                if text:
                    history_section.append(text)
            next_sib = next_sib.next_sibling
        break

# 输出历史板块内容
if history_section:
    for content in history_section:
        print(content)

说明

  • 方法一适合精准匹配指定区间,但依赖元素的属性/文本稳定性;
  • 方法二更适配维基百科的固定章节结构,容错性更高,优先推荐使用。

内容的提问来源于stack exchange,提问作者Zachariah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 10:25:35