如何用BeautifulSoup提取HTML中body标签的data-is-article属性值
用BeautifulSoup提取body标签的data-is-article属性值
你之前的代码返回None是因为用错了参数:class_是BeautifulSoup专门用来匹配HTML元素class属性的参数,不能用来匹配自定义的data-is-article属性。
下面是两种正确的实现方式:
方式一:通过attrs参数匹配属性并提取
from bs4 import BeautifulSoup html = '<body class="" data-is-article="story" data-new-gr-c-s-check-loaded="14.1094.0" data-gr-ext-installed=""></body>' soup = BeautifulSoup(html, 'html.parser') # 匹配存在data-is-article属性的body标签 body_tag = soup.find('body', attrs={'data-is-article': True}) if body_tag: article_type = body_tag['data-is-article'] print(article_type) # 输出结果:story
方式二:先定位body标签再提取属性(更简洁)
这种方式不需要提前匹配属性,直接获取目标属性值,用get()方法还能避免属性不存在时抛出异常:
from bs4 import BeautifulSoup html = '<body class="" data-is-article="story" data-new-gr-c-s-check-loaded="14.1094.0" data-gr-ext-installed=""></body>' soup = BeautifulSoup(html, 'html.parser') body_tag = soup.find('body') if body_tag: # get()方法在属性不存在时返回None,更安全 article_type = body_tag.get('data-is-article') print(article_type) # 输出结果:story # 如果确定body标签和目标属性一定存在,可简化为: article_type = soup.find('body')['data-is-article']
内容的提问来源于stack exchange,提问作者tphr
相关产品推荐
相关产品推荐

