如何使用Cheerio遍历<p>元素的所有子节点(含纯文本)
Cheerio获取
元素所有子节点(含纯文本节点)的解决方案
在遍历Cheerio中<p>元素的子节点时,使用find('*')或children('*')只能获取到<strong>这类元素节点,无法捕获纯文本节点。需要获取<p>下所有嵌套元素和纯文本的完整列表。
给定HTML结构:
<html> <head> </head> <body> <p>Hello, this is me - Daniel</p> <p><strong>Hello</strong>, this is me - Daniel</p> <p>Hello, <strong>this is me</strong> - Daniel</p> <p>Hello, this is me - <strong>Norbert</strong></p> <p><strong>Hello</strong>, this is me - <strong>Daniel</strong></p> </body> </html>
解决方案
Cheerio的children('*')和find('*')仅返回元素节点,要获取所有子节点(包括文本、注释等),需使用contents()方法。该方法会返回当前元素的所有子节点,之后可通过节点类型区分处理。
代码示例
const cheerio = require('cheerio'); const html = ` <html> <head> </head> <body> <p>Hello, this is me - Daniel</p> <p><strong>Hello</strong>, this is me - Daniel</p> <p>Hello, <strong>this is me</strong> - Daniel</p> <p>Hello, this is me - <strong>Norbert</strong></p> <p><strong>Hello</strong>, this is me - <strong>Daniel</strong></p> </body> </html> `; const $ = cheerio.load(html); // 遍历每个<p>元素,处理所有子节点 $('p').each((pIndex, pElement) => { console.log(`=== 第${pIndex + 1}个<p>的子节点 ===`); $(pElement).contents().each((nodeIndex, node) => { // 处理文本节点:过滤空白内容 if (node.type === 'text') { const cleanText = $(node).text().trim(); if (cleanText) { console.log(`文本节点:${cleanText}`); } } // 处理元素节点 else if (node.type === 'tag') { console.log(`元素节点 <${node.name}>:${$(node).text().trim()}`); } }); });
代码说明
$(pElement).contents():获取当前<p>下的所有子节点,包含文本、元素等类型。node.type:判断节点类型,text对应纯文本节点,tag对应HTML元素节点。trim():去除文本节点中的空白字符(换行、空格等),避免输出无意义的空文本。
结构化列表整理
如果需要将所有子节点整理为结构化数据列表,可参考以下代码:
const childNodeList = []; $('p').contents().each((_, node) => { if (node.type === 'text') { const text = $(node).text().trim(); text && childNodeList.push({ type: 'text', content: text }); } else if (node.type === 'tag') { childNodeList.push({ type: 'element', tagName: node.name, content: $(node).text().trim() }); } }); console.log('所有子节点列表:', childNodeList);
内容的提问来源于stack exchange,提问作者Norbert Hüthmayr
相关产品推荐
相关产品推荐

