You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用XPath从指定HTML结构中提取发布时间文本

XPath提取指定文本实现方案

待解析的目标HTML结构如下(原代码中div的class属性存在单双引号混用的笔误,HTML解析器会自动修正class值为postbodytop,不影响节点定位):

<div class="postbodytop">
  <a class="xxxxxxxxxxxxxxxx" href="xxxxxxxxxxxxxx">tonyd</a>
  "posted this 4 minutes ago "
  <span class="hidden-xs"> </span>  
</div>

提取完整文本posted this 4 minutes ago

目标文本是div节点下的直接文本节点,不属于a、span等子元素的内容,可按需选择表达式:

  • 基础定位表达式(取到原始文本节点,包含前后多余空格、包裹的双引号):
    //div[contains(@class,'postbodytop')]/text()[normalize-space()]
    
  • 格式化输出表达式(自动去除前后空白、剥离包裹的双引号,兼容所有XPath 1.0环境):
    normalize-space(translate(//div[contains(@class,'postbodytop')]/text()[normalize-space()], '"', ''))
    
    执行后直接返回无多余格式的posted this 4 minutes ago。

提取时间片段4 minutes

根据使用环境的XPath版本支持情况,可选两种写法:

  • 兼容XPath 1.0(无正则依赖,基于固定文本结构截取):
    substring-before(substring-after(normalize-space(translate(//div[contains(@class,'postbodytop')]/text()[normalize-space()], '"', '')), 'this '), ' ago')
    
  • XPath 2.0+环境(正则匹配,容错性更高,不受前后文本微调影响):
    replace(normalize-space(//div[contains(@class,'postbodytop')]/text()[normalize-space()]), '^.*?(\d+\s+minutes).*$', '$1')
    
    上述两个表达式执行后都会返回目标时间片段4 minutes。

内容的提问来源于stack exchange,提问作者Sam Rodriguez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 17:39:18