Cheerio返回不存在的匹配结果?网页爬取问题求助
问题:爬取时生成大量无效图片项
我是JavaScript新手,正在爬取一个结构混乱的网站,用Node.js结合Cheerio写的代码能获取任务标题、文本内容、图片及链接,但会额外返回大量{ type: 'image', src: undefined }的无效匹配项。
期望输出结构
{ "tasks": { "0": { "title": "Lorem ipsum dolor sit amet.", "contents": { "0": { "type": "text", "content": "Lorem .... " }, "1": { "type": "image", "src": "image.jpg" }, "2": { "type": "text", "content": "Yada yada yad", "links": { "0": { "linkText": "yada", "href": "someplace.html" }, "1": { "linkText": "yad", "href": "someotherplace.html" } } } } } } }
目标网页HTML
<div class="section" id="unit-tasks"> <h4 class="section__title"> Tasks <ul class="course--tasks-check-list"> <li class="task-item"> <div class="task-item__anchor"> <h3 class="task-item__title"> Lorem ipsum dolor sit amet.<i class="fa"></i><span class="badge">Easy Task</span> </h3> <div class="task-item__description"> <p>Lorem ipsum dolor, sit amet consectetur adipisicing elit.</p> <p>Unde velit tenetur quia architecto tempora molestiae?</p> <div class="image-item"> <img src="image.jpg" /> </div> <p>Lorem ipsum <a href="someplace.html">dolor, sit amet</a> consectetur <a href="elsewhere.html"> adipisicing elit.</a> Quis, facilis?</p> </div> </div> </li> <li class="task-item"> <div class="task-item__anchor"> <h3 class="task-item__title"> Explicabo mollitia molestias illum maxime.<i class="fa"></i><span class="badge">Medium task</span> </h3> <div class="task-item__description"> <p>Lorem ipsum dolor, sit amet consectetur adipisicing elit.</p> <div class="btn-toolbar"> <a class="btn btn-secondary btn-sm" href="sendhome.html">Lorem!</a> </div> </div> </div> </li> </ul> </h4> </div>
当前代码
const cheerio = require('cheerio') const fs = require('fs'); const chapTasks = {} function getTasks(){ let taskCount = 0 if ($('.task-item').length) { $('.task-item').each( function() { chapTasks[taskCount] = {} chapTasks[taskCount]['contents'] = {} taskTitle = $(this).find('.task-item__title').html() taskTitle = taskTitle.slice(0, taskTitle.indexOf('<')).trim() chapTasks[taskCount]['title'] = taskTitle let contentCount = 0 $(this).find('.task-item__description').children().each( function() { let tag = $(this).get(0).tagName if (tag == 'p') { chapTasks[taskCount]['contents'][contentCount] = { type : 'text', content : $(this).text().replace(/\s+/g, ' ').trim() } if ($(this).find('a').length){ linkCount = 0 chapTasks[taskCount]['contents'][contentCount]['links'] = {} $(this).find('a').each( function() { chapTasks[taskCount]['contents'][contentCount]['links'][linkCount] = { 'linkText' : $(this).text(), 'href' : $(this).attr('href') } linkCount += 1 }) } contentCount += 1 } if ($(this).find('div.image-item > img')) { chapTasks[taskCount]['contents'][contentCount] = { type : 'image', 'src' : $(this).find('img').attr('src') } contentCount += 1 } }) taskCount += 1 }) } console.dir(chapTasks, { depth: null }) } getTasks()
原因排查
- 图片判断逻辑失效:
if ($(this).find('div.image-item > img'))永远为真,因为Cheerio查询返回的是对象,即使没有匹配到元素也不会返回null或undefined。所以遍历到所有非<p>的子元素(比如第二个任务里的.btn-toolbar)时,都会执行这个分支,找不到img就会生成src: undefined的无效项。 - 未校验src有效性:即使找到img元素,也没有判断
src属性是否存在,若遇到无src的img也会生成无效数据。
修复方案
1. 修正图片匹配逻辑
把原来的图片判断代码替换为:
// 只处理带有image-item类的容器,且确保img存在且有src if ($(this).hasClass('image-item')) { const img = $(this).find('img'); const src = img.attr('src'); if (src) { chapTasks[taskCount]['contents'][contentCount] = { type: 'image', src: src.trim() }; contentCount += 1; } }
2. 优化标题提取方式
原来的html().slice(0, indexOf('<'))依赖HTML结构的顺序,改用文本节点提取更可靠:
// 替换原来的标题提取代码 const taskTitle = $(this).find('.task-item__title').contents().first().text().trim(); chapTasks[taskCount]['title'] = taskTitle;
3. 完整修复后的代码
const cheerio = require('cheerio') const fs = require('fs'); const chapTasks = { tasks: {} }; // 对齐期望的外层tasks结构 function getTasks(){ let taskCount = 0 if ($('.task-item').length) { $('.task-item').each( function() { chapTasks.tasks[taskCount] = { contents: {} }; // 优化标题提取 const taskTitle = $(this).find('.task-item__title').contents().first().text().trim(); chapTasks.tasks[taskCount]['title'] = taskTitle; let contentCount = 0 $(this).find('.task-item__description').children().each( function() { const tag = $(this).get(0).tagName.toLowerCase(); // 统一小写避免大小写问题 if (tag === 'p') { const textContent = $(this).text().replace(/\s+/g, ' ').trim(); chapTasks.tasks[taskCount]['contents'][contentCount] = { type : 'text', content : textContent }; // 处理链接 const links = $(this).find('a'); if (links.length){ let linkCount = 0; chapTasks.tasks[taskCount]['contents'][contentCount]['links'] = {}; links.each( function() { const linkText = $(this).text().trim(); const href = $(this).attr('href'); if (linkText && href) { // 校验链接有效性 chapTasks.tasks[taskCount]['contents'][contentCount]['links'][linkCount] = { linkText, href: href.trim() }; linkCount += 1; } }); // 如果没有有效链接,删除links属性 if (linkCount === 0) { delete chapTasks.tasks[taskCount]['contents'][contentCount]['links']; } } contentCount += 1; } // 修正图片判断逻辑 if ($(this).hasClass('image-item')) { const img = $(this).find('img'); const src = img.attr('src'); if (src) { chapTasks.tasks[taskCount]['contents'][contentCount] = { type: 'image', src: src.trim() }; contentCount += 1; } } }) taskCount += 1; }) } console.dir(chapTasks, { depth: null }); } getTasks();
更合理的爬取思路
- 用数组存储列表数据:把
tasks、contents、links改成数组,比如tasks: [],用push方法添加元素,避免手动维护计数变量,代码更简洁易读。 - 封装复用函数:把提取标题、处理段落、处理图片的逻辑拆成小函数,比如
extractTaskTitle($el)、processParagraph($el),降低代码耦合度。 - 精准选择器:直接用
.task-item__description > p, .task-item__description > .image-item定位需要处理的元素,减少不必要的遍历。 - 严格数据校验:对所有提取的字段(比如src、href、文本内容)做非空校验,避免生成无效数据。
内容的提问来源于stack exchange,提问作者Joao Mendonca
相关产品推荐
相关产品推荐

