You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用CSS选择器或XPath从特殊结构HTML中提取资格文本?

精准提取资格信息的CSS选择器与XPath方案

我有一个网站,其内部包含如下HTML结构:

<div class="ui-rectframe">
    <p class="ui-li-desc"></p>
    <h4 class="ui-li-heading">Qualifications</h4>
    MBBS (University of Singapore, Singapore) 1978
    <br>
    MCFP (Family Med) (College of Family Physicians, Singapore) 1984
    <br>
    Dip Geriatric Med (NUS, Singapore) 2012
    <br>
    GDPM (NUS, Singapore) 2015
    <br>
    <h4 class="ui-li-heading">Type of first registration / date</h4>
    Full Registration (14/06/1979)<br>
    <h4 class="ui-li-heading">Type of current registration / date</h4>
    Full Registration (14/06/1979)<br>
    <h4 class="ui-li-heading">Practising Certificate Start Date</h4>
    01/01/2022<br>
    <h4 class="ui-li-heading">Practising Certificate End Date</h4>
    31/12/2023<br>
    <p></p><br>
</div>

我需要提取资格信息列表:[ 'MBBS (University of Singapore, Singapore) 1978', 'MCFP (Family Med) (College of Family Physicians, Singapore) 1984', 'Dip Geriatric Med (NUS, Singapore) 2012', 'GDPM (NUS, Singapore) 2015' ]
目前我能提取该父div内的所有文本,但无法将资格信息与其他注册类信息区分开,请问如何使用CSS选择器或XPath实现精准提取?


XPath 方案

XPath可直接定位目标文本节点,精准筛选出"Qualifications"标题后、下一个h4标签前的内容:

核心表达式

//div[@class='ui-rectframe']/h4[text()='Qualifications']/following-sibling::text()[following-sibling::h4[1][text()='Type of first registration / date']]

逻辑说明

  1. //div[@class='ui-rectframe']:定位到包含所有信息的父容器
  2. /h4[text()='Qualifications']:找到标记资格信息的标题标签
  3. /following-sibling::text():选取该标题之后的所有同级文本节点
  4. [following-sibling::h4[1][text()='Type of first registration / date']]:过滤出仅在下一个注册类标题前的文本节点

提取后只需清洗掉空白字符、过滤空文本,就能得到目标资格列表。


CSS选择器方案

CSS选择器无法直接选取文本节点,需结合代码逻辑完成筛选:

步骤1:定位父容器

用CSS选择器定位目标容器:

div.ui-rectframe

步骤2:结合代码过滤内容

以JavaScript为例,通过DOM遍历精准提取文本节点:

// 找到资格信息的标题标签
const qualificationsTitle = document.querySelector('div.ui-rectframe h4.ui-li-heading:nth-of-type(1)');
// 找到下一个注册类标题标签作为边界
const nextSectionTitle = qualificationsTitle.nextElementSibling;

const qualificationList = [];
let currentNode = qualificationsTitle.nextSibling;

// 遍历两个标题之间的所有节点
while (currentNode && currentNode !== nextSectionTitle) {
  // 只保留非空的文本节点
  if (currentNode.nodeType === Node.TEXT_NODE && currentNode.textContent.trim()) {
    qualificationList.push(currentNode.textContent.trim());
  }
  currentNode = currentNode.nextSibling;
}
// qualificationList 即为目标资格信息列表

内容的提问来源于stack exchange,提问作者Artsiom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 12:01:04