You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取指定网页中Barry Kripke的完整内容?

解决Fandom页面内容抓取问题

问题根源

你之前的代码仅提取了页面第一个<p>标签的内容,自然只能拿到第一段;提取链接时也仅在该段落内查找,所以结果有限。Fandom(基于MediaWiki)的正文内容都集中在div.mw-parser-output容器中,这是排除页眉、页脚等无关元素的核心定位点。

修改后的完整代码

import urllib3
from bs4 import BeautifulSoup

url = 'https://bigbangtheory.fandom.com/wiki/Barry_Kripke'
http = urllib3.PoolManager()
r = http.request('GET', url)

if r.status == 200:
    page = r.data
    print('页面请求成功,状态码:%d,页面大小:%d' % (r.status, len(page)))
else:
    print('请求失败,状态码:%s' % r.status)
    exit()

soup = BeautifulSoup(page, 'html.parser')

# 定位核心内容容器,自动排除无关元素
main_content = soup.find('div', class_='mw-parser-output')

if not main_content:
    print('未找到主要内容容器')
    exit()

# 提取所有正文元素(段落、标题、列表等)
all_content = []
for element in main_content.find_all(['p', 'h2', 'h3', 'ul', 'ol']):
    # 清理标题中的编辑链接
    if element.name in ['h2', 'h3']:
        edit_section = element.find('span', class_='mw-editsection')
        if edit_section:
            edit_section.decompose()
        all_content.append(f'**{element.get_text(strip=True)}**')
    else:
        all_content.append(element.get_text(strip=True))

# 输出完整正文内容
print('\n'.join(all_content))

# 提取容器内所有有效链接
print('\n=== 页面有效链接 ===')
for link in main_content.find_all('a', href=True):
    # 过滤编辑页、锚点等无效链接
    if not link['href'].startswith(('/wiki/Special:', '#')):
        print(f'文本:{link.get_text(strip=True)},链接:{link["href"]}')

关键改进点

  • 精准定位内容区:通过div.mw-parser-output直接锁定角色相关的全部正文,无需手动过滤页眉页脚。
  • 覆盖多类型元素:除了段落<p>,还包含标题、列表等元素,确保内容完整。
  • 清理冗余内容:移除标题旁的“编辑”链接,避免干扰阅读。
  • 过滤无效链接:排除内部编辑页面、锚点链接,只保留有价值的内容链接。

内容的提问来源于stack exchange,提问作者WMD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 00:18:31