如何用X-ray将非嵌套HTML结构解析为对象数组?
问题描述
需要用X-ray采集非嵌套结构的页面数据,页面HTML结构如下:
<h1>Page title</h1> <article> <h2 id="first">Title 1</h2> <h3>Subtitle 1</h3> <ul> <li>Element 1 <li>Element 2 <li>Element 3 </ul> <h2 id="second">Title 2</h2> <h3>Subtitle 2</h3> <h2 id="third">Title 3</h2> <h3>Subtitle 3</h3> <ul> <li>Element 1 <li>Element 2 <li>Element 3 </ul> </article>
目标是生成如下结构的数据:
type Result = { pageTitle: string, sections: [{ subtitle?: string, elements?: string[] }], }
期望输出示例:
{ "pageTitle": "Page title", "sections": [ { "subtitle": "Subtitle 1", "elements": ["Element 1", "Element 2", "Element 3"] }, { "subtitle": "Subtitle 2", "elements": [] }, { "subtitle": "Subtitle 3", "elements": ["Element 1", "Element 2", "Element 3"] } ] }
尝试过两种写法均失败:
- 以
article为根遍历,因只有一个article标签无法生成多个section:
xray(url, { pageTitle: "h1 | trim", sections: xray("article", [{ subtitle: "h3", elements: ['h3 ~ ul li'] }]) })
- 以
h2为根但作用域限制在h2内部,无法获取兄弟元素:
xray(url, { pageTitle: "h1 | trim", sections: xray("h2", [{ subtitle: "h3", elements: ['h3 ~ ul li'] }]) })
解决方案
使用CSS相邻兄弟选择器+和通用兄弟选择器~,结合X-ray的遍历功能,就能正确获取每个h2对应的后续内容:
xray(url, { pageTitle: "h1 | trim", sections: xray("h2", [{ // 选择当前h2的直接下一个兄弟h3元素 subtitle: "+ h3 | trim", // 选择当前h2之后的h3对应的ul下的所有li elements: "+ h3 ~ ul li | trim" }]) })
原理说明
- 以
h2作为遍历节点,每个h2对应一个section,自然生成数组结构。 + h3表示当前h2的直接下一个兄弟元素h3,确保每个章节的副标题对应正确。+ h3 ~ ul li表示当前h2之后的h3元素后面的ul中的li,如果该章节没有ul,X-ray会自动返回空数组,符合需求。- 保留
trim过滤器去除文本多余空格,保证数据整洁。
内容的提问来源于stack exchange,提问作者IllidariCat
相关产品推荐
相关产品推荐

