使用await后仍无法同步获取数据?网页爬取问题求助
问题原因
数据错乱的核心问题在于:
processhtmldata是异步函数,但你用forEach遍历调用它时,forEach不会等待异步函数执行完成,所有processhtmldata会并行启动。- 每个
processhtmldata内部的saveimagetosystem是异步操作,不同URL的图片保存耗时不同,先完成的会先修改details对象,导致最终数据顺序和原始URL顺序不一致。
解决方案
用 for...of 循环替代 forEach,强制异步操作按顺序执行,确保每个URL的处理完成后再进行下一个,这样 details 中的数据就会严格按照原始URL的顺序存储。
同时修正代码中的几个细节问题:
- 变量名拼写错误:
imeg→image getElementsByTagName永远返回HTMLCollection(不会是null),判断元素是否存在应该用.length > 0而非!= null
修改后的完整代码
1. 获取HTML数据部分(无改动)
async function getdata(value){ let data = ""; let txtPromise; if(value.includes('https')){ txtPromise = await fetch('https://api.codetabs.com/v1/proxy?quest='+value); if (txtPromise.ok) { data = await txtPromise.text(); } else{ data = "not found"; } } else data = "invalid/null url"; return data; } let urls = ["https://.......","https://......","https://...."]; const responses = []; // 补充声明responses变量 for(var i=0; i<urls.length; i++){ responses.push(getdata(urls[i])); }
2. 解析处理数据部分(关键修改)
function saveimagetosystem(url){ return $.ajax({ url: '/server.php', type: 'POST', data:{"input":"save image","url":url}, cache:false, }) } async function processhtmldata(html_data,index){ if(!html_data.includes("not found") && !html_data.includes("invalid/null url")){ var parser = new DOMParser(); var htmldoc = parser.parseFromString(html_data, "text/html"); var title = htmldoc.querySelector("title").innerText; details["title"+" "+index]=title.trim(); // 修正:判断img元素是否存在 if(htmldoc.getElementsByTagName("img").length > 0){ var url = htmldoc.getElementsByTagName("img")[0].src.trim(); var response = await saveimagetosystem('https://api.codetabs.com/v1/proxy?quest='+url); // 修正拼写:imeg → image details["image"+" "+index] = response; // 修正:判断p元素是否存在 if(htmldoc.getElementsByTagName("p").length > 0){ for(var l=0; l<htmldoc.getElementsByTagName("p").length-1; l++){ details["body"+" "+index+" "+l] = htmldoc.getElementsByTagName("p")[l].innerText.trim(); } } } } } // 关键修改:用for...of循环保证顺序执行 Promise.all(responses) .then(async (htmlfiles) => { for (let sl_no = 0; sl_no < htmlfiles.length; sl_no++) { await processhtmldata(htmlfiles[sl_no], sl_no); } // 处理完成后可以在这里访问正确顺序的details对象 console.log(details); }) .catch(err => console.error('处理失败:', err));
额外优化(并行处理+保序)
如果希望保持并行处理以提高效率,同时保证数据顺序正确,可以收集所有processhtmldata的Promise,再用Promise.all等待全部完成。这种方式下,虽然处理过程是并行的,但每个processhtmldata会把数据写入对应index的字段,最终details的顺序依然和原始URL一致:
Promise.all(responses) .then(async (htmlfiles) => { const processPromises = htmlfiles.map((file, sl_no) => processhtmldata(file, sl_no)); await Promise.all(processPromises); console.log(details); })
内容的提问来源于stack exchange,提问作者user20643052
相关产品推荐
相关产品推荐

