You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取HTML中body标签的data-is-article属性值

用BeautifulSoup提取body标签的data-is-article属性值

你之前的代码返回None是因为用错了参数:class_是BeautifulSoup专门用来匹配HTML元素class属性的参数,不能用来匹配自定义的data-is-article属性。

下面是两种正确的实现方式:

方式一:通过attrs参数匹配属性并提取

from bs4 import BeautifulSoup

html = '<body class="" data-is-article="story" data-new-gr-c-s-check-loaded="14.1094.0" data-gr-ext-installed=""></body>'
soup = BeautifulSoup(html, 'html.parser')

# 匹配存在data-is-article属性的body标签
body_tag = soup.find('body', attrs={'data-is-article': True})
if body_tag:
    article_type = body_tag['data-is-article']
    print(article_type)  # 输出结果:story

方式二:先定位body标签再提取属性(更简洁)

这种方式不需要提前匹配属性,直接获取目标属性值,用get()方法还能避免属性不存在时抛出异常:

from bs4 import BeautifulSoup

html = '<body class="" data-is-article="story" data-new-gr-c-s-check-loaded="14.1094.0" data-gr-ext-installed=""></body>'
soup = BeautifulSoup(html, 'html.parser')

body_tag = soup.find('body')
if body_tag:
    # get()方法在属性不存在时返回None,更安全
    article_type = body_tag.get('data-is-article')
    print(article_type)  # 输出结果:story

# 如果确定body标签和目标属性一定存在,可简化为:
article_type = soup.find('body')['data-is-article']

内容的提问来源于stack exchange,提问作者tphr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 10:55:19