You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Sheets ImportXML使用指南:如何提取指定文本标签后的对应内容

提取指定文本标签后内容的解决方案

不推荐直接使用正则匹配HTML结构,页面的空格、不可见字符、标签属性微调都会导致正则失效,优先使用成熟HTML解析库的节点关系查询能力实现需求,稳定性更高。
核心实现逻辑:先定位到包含指定文本「Publication date」的标签,再通过节点的相邻、父子层级关系找到后续的目标内容节点。

示例1实现代码

通用XPath写法(适配绝大多数爬虫工具、解析库)

//div[contains(@class, 'rpi-attribute-label')]/span[text()='Publication date']/../../div[contains(@class, 'rpi-attribute-value')]/span/text()

定位逻辑:

  • 先找到class含rpi-attribute-label的div下,文本值为「Publication date」的span节点
  • 向上回溯两层到属性组的根容器rpi-attribute-content
  • 向下查找class含rpi-attribute-value的div下的span文本,即为目标日期

Python BeautifulSoup 示例

from bs4 import BeautifulSoup

# 传入你抓取到的HTML文本
soup = BeautifulSoup(html_content, 'lxml')
# 定位Publication date标签
label_node = soup.find('span', string='Publication date')
# 向上找到属性组容器
attr_group = label_node.parent.parent
# 提取目标日期
target_date = attr_group.find('div', class_='rpi-attribute-value').span.text.strip()
print(target_date) # 输出:May 11, 2021

示例2实现代码

通用XPath写法

//span[contains(@class, 'a-text-bold') and contains(text(), 'Publication date')]/following-sibling::span[1]/text()

定位逻辑:

  • 先找到class含a-text-bold、文本包含「Publication date」的span节点
  • 直接取该节点后第一个同级span的文本即可

Python BeautifulSoup 示例

from bs4 import BeautifulSoup

soup = BeautifulSoup(html_content, 'lxml')
# 定位Publication date加粗标签
label_node = soup.find('span', class_='a-text-bold', string=lambda s: 'Publication date' in s)
# 取后面第一个同级span的内容
target_date = label_node.find_next_sibling('span').text.strip()
print(target_date) # 输出:September 7, 2021

兜底替代方案

仅当页面结构极不稳定、节点定位逻辑频繁失效时,可考虑使用正则做模糊匹配,示例规则如下:

Publication date[\s\S]*?([A-Z][a-z]+ \d{1,2}, \d{4})

该方案容错率低,仅作为兜底使用,优先级远低于HTML节点解析方案

内容的提问来源于stack exchange,提问作者RandX

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 03:36:03