如何使用XPath从指定HTML结构中提取发布时间文本
XPath提取指定文本实现方案
待解析的目标HTML结构如下(原代码中div的class属性存在单双引号混用的笔误,HTML解析器会自动修正class值为postbodytop,不影响节点定位):
<div class="postbodytop"> <a class="xxxxxxxxxxxxxxxx" href="xxxxxxxxxxxxxx">tonyd</a> "posted this 4 minutes ago " <span class="hidden-xs"> </span> </div>
提取完整文本posted this 4 minutes ago
目标文本是div节点下的直接文本节点,不属于a、span等子元素的内容,可按需选择表达式:
- 基础定位表达式(取到原始文本节点,包含前后多余空格、包裹的双引号):
//div[contains(@class,'postbodytop')]/text()[normalize-space()] - 格式化输出表达式(自动去除前后空白、剥离包裹的双引号,兼容所有XPath 1.0环境):
执行后直接返回无多余格式的normalize-space(translate(//div[contains(@class,'postbodytop')]/text()[normalize-space()], '"', ''))posted this 4 minutes ago。
提取时间片段4 minutes
根据使用环境的XPath版本支持情况,可选两种写法:
- 兼容XPath 1.0(无正则依赖,基于固定文本结构截取):
substring-before(substring-after(normalize-space(translate(//div[contains(@class,'postbodytop')]/text()[normalize-space()], '"', '')), 'this '), ' ago') - XPath 2.0+环境(正则匹配,容错性更高,不受前后文本微调影响):
上述两个表达式执行后都会返回目标时间片段replace(normalize-space(//div[contains(@class,'postbodytop')]/text()[normalize-space()]), '^.*?(\d+\s+minutes).*$', '$1')4 minutes。
内容的提问来源于stack exchange,提问作者Sam Rodriguez
相关产品推荐
相关产品推荐

