You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取带有data-wipe-name属性的h2标签文本内容?

提取方案汇总

下面是不同常用爬虫工具下的实现方式,均针对你给出的<h2 data-wipe-name='Titel'>Company name</h2>标签结构:

1. Python + BeautifulSoup4(新手友好,最常用)

首先安装依赖:pip install beautifulsoup4 lxml
示例代码:

from bs4 import BeautifulSoup

# 假设你已经拿到的页面HTML源码存在html_content变量里
soup = BeautifulSoup(html_content, 'lxml')
# 按自定义属性定位目标h2标签
target_h2 = soup.find('h2', {'data-wipe-name': 'Titel'})
# 提取文本
if target_h2:
    company_name = target_h2.get_text(strip=True)
    print(company_name) # 输出结果:Company name

2. Python + lxml(性能更高)

安装依赖:pip install lxml
示例代码:

from lxml import etree

html = etree.HTML(html_content)
# xpath语法定位标签并提取文本
company_name = html.xpath("//h2[@data-wipe-name='Titel']/text()")[0].strip()
print(company_name)

3. 正则表达式(仅作为固定结构下的应急方案,不推荐通用场景使用)

如果页面结构完全固定没有变动,也可以用正则快速匹配:

import re

match_res = re.search(r"<h2 data-wipe-name='Titel'>(.*?)</h2>", html_content)
if match_res:
    company_name = match_res.group(1).strip()
    print(company_name)

注意事项

  • 定位标签优先用唯一属性,你给出的data-wipe-name='Titel'属于自定义属性,只要页面内该属性唯一,以上定位方式不会出错
  • 代码里的strip()方法用于去除文本前后的空格、换行符,不需要可以删掉
  • 如果实际页面中该h2标签还有其他唯一类名、id,可以叠加到定位条件里,进一步提升匹配准确率

内容的提问来源于stack exchange,提问作者Devindi Siwurathna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 21:09:03