Axios遍历响应重复返回首个映射对象结果的修复方案咨询
问题分析与修复方案
问题根源
- 全局共享的
articles数组:所有请求都往同一个全局数组中追加内容,且每次返回整个数组。由于异步请求的执行顺序不确定,后续请求的返回结果会包含前面所有报纸的内容,最终导致嵌套数组中每个元素都是累积的重复结果。 - 全局
tempUrls未重置:如果tempUrls是全局变量,处理完第一个报纸后,它会存储该报纸的所有URL。后续报纸的内容会被错误过滤(或因全局状态干扰导致爬取逻辑异常),最终只能返回首个报纸的结果。
修复代码
关键修改点
- 移除全局
articles变量,改为每个请求独立处理自身结果 - 将
tempUrls改为局部变量,确保单报纸爬取时的去重逻辑独立 - 函数参数化依赖(如
exceptions),避免隐式全局依赖
const newspapers= [{ "name": "CNN", "address": "https://edition.cnn.com/specials/world/cnn-climate", "base": "https://edition.cnn.com" }, { "name": "The Guardian", "address": "https://www.theguardian.com/environment/climate-crisis", "base": "https://www.theguardian.com" }]; // 假设exceptions是预定义的需要排除的链接关键词数组 const exceptions = ["twitter.com", "facebook.com"]; // 重构storeData:接收tempUrls和exceptions作为参数 function storeData(element, base, name, tempUrls) { const results = []; element.find("style").remove(); const title = element.text(); const urlRaw = element.attr("href"); const url = urlRaw.includes("www") || urlRaw.includes("http") ? urlRaw : base + urlRaw; // 仅对当前报纸的URL去重 if (tempUrls.indexOf(url) === -1) { if (!exceptions.some((el) => url.toLowerCase().includes(el))) { tempUrls.push(url); const imageElement = element.find("img"); if (imageElement.length > 0) { results.push({ title, url, source: name, imgUrl: getImageFromElement(imageElement), }); } else { results.push({ title, url: url, source: name, }); } } } return results; } // 重构getElementsCheerio:创建局部tempUrls function getElementsCheerio(html, base, name, searchterms) { const $ = cheerio.load(html); const termsAlso = searchterms.also; const termsOnly = searchterms.only; const concatInfo = []; const tempUrls = []; // 局部变量,仅当前报纸爬取时使用 termsAlso.forEach((term) => { $(`a:contains("climate"):contains(${term})`).each(function () { const tempData = storeData($(this), base, name, tempUrls); tempData.forEach(el => concatInfo.push(el)); }); }); termsOnly.forEach((term) => { $(`a:contains(${term})`).each(function () { const tempData = storeData($(this), base, name, tempUrls); tempData.forEach(el => concatInfo.push(el)); }); }); return concatInfo; } // 重构API路由:移除全局articles,返回单报纸独立结果 app.get("/news", (req, res) => { const query = checkForQuery(req); const wordsToSearch = query ? verifyQuery(query) : ""; Promise.all( newspapers.map(({ name, address, base }) => axios .get(address, { headers: { "Accept-Encoding": "gzip,deflate,compress" }, }) .then((axiosRes) => { const html = axiosRes.data; console.log({ name, address, base }); const scrappedElements = getElementsCheerio( html, base, name, wordsToSearch ); // 返回当前报纸的独立结果,而非全局累积数组 return scrappedElements; }) ) ).then((allResults) => { // 可选:如果需要一维数组而非嵌套数组,使用.flat() // res.json(allResults.flat()); res.json(allResults); }); });
修复说明
- 独立结果处理:每个axios请求爬取完成后,直接返回当前报纸的文章数组,不再依赖全局变量,确保每个请求的结果独立。
- 局部去重逻辑:
tempUrls被移到getElementsCheerio内部,仅对当前报纸的URL去重,不会干扰其他报纸的爬取。 - 语义化遍历:用
forEach替代map执行数组追加操作,map的语义是转换数组,而非执行副作用。
内容的提问来源于stack exchange,提问作者Bella
相关产品推荐
相关产品推荐

