如何用Beautiful Soup抓取维基百科特定板块内的多元素
问题描述
我是从C#转Python的新手,想要抓取维基百科微软页面中历史板块的全部文本内容。已知Beautiful Soup不支持XPath,该板块的第一个元素为:
<div role="note" class="hatnote navigation-not-searchable">Main article: <a href="/wiki/History_of_Microsoft" title="History of Microsoft">History of Microsoft</a></div>
最后一个元素为:
<p>On January 18, 2022, Microsoft announced the acquisition of American video game developer and <a href="/wiki/Holding_company" title="Holding company">holding company</a> <a href="/wiki/Activision_Blizzard" title="Activision Blizzard">Activision Blizzard</a> in an all-cash deal worth $68.7 billion.<sup id="cite_ref-:0_150-0" class="reference"><a href="#cite_note-:0-150">[150]</a></sup> Activision Blizzard is best known for producing franchises, including but not limited to <i><a href="/wiki/Warcraft" title="Warcraft">Warcraft</a></i>, <i><a href="/wiki/Diablo_(series)" title="Diablo (series)">Diablo</a></i>, <i><a href="/wiki/Call_of_Duty" title="Call of Duty">Call of Duty</a></i>, <i><a href="/wiki/StarCraft" title="StarCraft">StarCraft</a></i>, <i><a href="/wiki/Candy_Crush_Saga" title="Candy Crush Saga">Candy Crush Saga</a></i>, <i><a href="/wiki/Crash_Bandicoot" title="Crash Bandicoot">Crash Bandicoot</a></i>, <i><a href="/wiki/Spyro" title="Spyro">Spyro the Dragon</a></i>, <i><a href="/wiki/Skylanders" title="Skylanders">Skylanders</a></i>, and <i><a href="/wiki/Overwatch_(video_game)" title="Overwatch (video game)">Overwatch</a></i>.<sup id="cite_ref-151" class="reference"><a href="#cite_note-151">[151]</a></sup> Activision and Microsoft each released statements saying the acquisition was to benefit their businesses in the <a href="/wiki/Metaverse" title="Metaverse">metaverse</a>, many saw Microsoft's acquisition of video game studios as an attempt to compete against <a href="/wiki/Meta_Platforms" title="Meta Platforms">Meta Platforms</a>, with <a href="/wiki/TheStreet" title="TheStreet">TheStreet</a> referring to Microsoft wanting to become "the <a href="/wiki/The_Walt_Disney_Company" title="The Walt Disney Company">Disney</a> of the metaverse".<sup id="cite_ref-152" class="reference"><a href="#cite_note-152">[152]</a></sup><sup id="cite_ref-153" class="reference"><a href="#cite_note-153">[153]</a></sup> Microsoft has not released statements regarding Activision's recent legal controversies regarding employee abuse, but reports have alleged that Activision CEO <a href="/wiki/Bobby_Kotick" title="Bobby Kotick">Bobby Kotick</a>, a major target of the controversy, will leave the company after the acquisition is finalized.<sup id="cite_ref-154" class="reference"><a href="#cite_note-154">[154]</a></sup> The deal is expected to close in 2023 followed by a review from the <a href="/wiki/US_Federal_Trade_Commission" class="mw-redirect" title="US Federal Trade Commission">US Federal Trade Commission</a>.<sup id="cite_ref-155" class="reference"><a href="#cite_note-155">[155]</a></sup><sup id="cite_ref-:0_150-1" class="reference"><a href="#cite_note-:0-150">[150]</a></sup></p>
请问如何获取这两个元素之间的所有信息?我已编写初始代码如下:
from bs4 import BeautifulSoup import requests url = 'https://en.wikipedia.org/wiki/Microsoft' # call get method to request that page page = requests.get(url) soup = BeautifulSoup(page.text, "html.parser")
解决方案
你可以通过两种方式实现需求,一种是精准定位起始/结束元素并遍历中间节点,另一种是直接定位整个历史章节(更稳定)。
方法一:基于起始/结束元素抓取
通过属性和文本特征定位目标元素,再遍历两者之间的所有兄弟节点:
from bs4 import BeautifulSoup import requests url = 'https://en.wikipedia.org/wiki/Microsoft' page = requests.get(url) soup = BeautifulSoup(page.text, "html.parser") # 定位起始元素:通过class和role组合匹配 start_element = soup.find('div', class_='hatnote navigation-not-searchable', role='note') # 定位结束元素:通过段落中的标志性文本匹配 end_element = soup.find('p', string=lambda text: text and 'On January 18, 2022, Microsoft announced the acquisition' in text) # 收集中间所有有效文本 history_content = [] current_element = start_element.next_sibling while current_element != end_element: # 过滤空节点和无文本的元素 if current_element and hasattr(current_element, 'get_text'): text = current_element.get_text(strip=True) if text: history_content.append(text) current_element = current_element.next_sibling # 添加结束元素的文本 if end_element: history_content.append(end_element.get_text(strip=True)) # 输出结果 for content in history_content: print(content)
方法二:直接定位历史章节(推荐)
维基百科的章节结构固定,通过标题History定位整个板块,避免单个元素变动导致失效:
from bs4 import BeautifulSoup import requests url = 'https://en.wikipedia.org/wiki/Microsoft' page = requests.get(url) soup = BeautifulSoup(page.text, "html.parser") history_section = [] # 遍历所有h2标题,找到History章节 for h2 in soup.find_all('h2'): if h2.get_text(strip=True) == 'History': # 收集h2之后的所有兄弟元素,直到下一个h2章节 next_sib = h2.next_sibling while next_sib and next_sib.name != 'h2': if next_sib and hasattr(next_sib, 'get_text'): text = next_sib.get_text(strip=True) if text: history_section.append(text) next_sib = next_sib.next_sibling break # 输出历史板块内容 if history_section: for content in history_section: print(content)
说明
- 方法一适合精准匹配指定区间,但依赖元素的属性/文本稳定性;
- 方法二更适配维基百科的固定章节结构,容错性更高,优先推荐使用。
内容的提问来源于stack exchange,提问作者Zachariah
相关产品推荐
相关产品推荐

