You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup匹配含部分值的data属性元素

正确匹配含指定字符串的data-title属性元素

你之前用"Subscribers"*的写法不符合BeautifulSoup的语法规则,以下是几种正确的实现方式:

方法1:正则表达式匹配

借助re模块,用正则匹配包含"Subscribers"的data-title属性值:

import re
from bs4 import BeautifulSoup

# 替换为你的实际HTML内容
html = '<div data-title="Subscribers: 4,471"></div>'
soup = BeautifulSoup(html, "html.parser")

# 匹配所有data-title包含Subscribers的div
target_divs = soup.find_all("div", {"data-title": re.compile(r"Subscribers")})

方法2:Lambda表达式筛选

通过lambda表达式检查元素是否有data-title属性,且属性值包含目标字符串:

from bs4 import BeautifulSoup

html = '<div data-title="Subscribers: 4,471"></div>'
soup = BeautifulSoup(html, "html.parser")

target_divs = soup.find_all(
    "div", 
    lambda tag: tag.has_attr("data-title") and "Subscribers" in tag["data-title"]
)

方法3:CSS选择器(推荐)

使用CSS选择器的*=语法,表示属性值包含指定字符串,写法更简洁:

from bs4 import BeautifulSoup

html = '<div data-title="Subscribers: 4,471"></div>'
soup = BeautifulSoup(html, "html.parser")

target_divs = soup.select('div[data-title*="Subscribers"]')

额外:提取订阅者数值

如果需要从匹配到的元素中提取具体的订阅者数量,可以这样处理:

if target_divs:
    title_text = target_divs[0]["data-title"]
    # 分割出数值部分,清理格式后转成整数
    subscriber_count = int(title_text.split(":")[1].strip().replace(",", ""))
    print(subscriber_count)  # 输出:4471

内容的提问来源于stack exchange,提问作者ethicnology

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 22:24:55