You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用await后仍无法同步获取数据?网页爬取问题求助

问题原因

数据错乱的核心问题在于:

  • processhtmldata 是异步函数,但你用 forEach 遍历调用它时,forEach 不会等待异步函数执行完成,所有 processhtmldata 会并行启动。
  • 每个 processhtmldata 内部的 saveimagetosystem 是异步操作,不同URL的图片保存耗时不同,先完成的会先修改 details 对象,导致最终数据顺序和原始URL顺序不一致。
解决方案

用 for...of 循环替代 forEach,强制异步操作按顺序执行,确保每个URL的处理完成后再进行下一个,这样 details 中的数据就会严格按照原始URL的顺序存储。

同时修正代码中的几个细节问题:

  1. 变量名拼写错误:imeg → image
  2. getElementsByTagName 永远返回 HTMLCollection(不会是 null),判断元素是否存在应该用 .length > 0 而非 != null
修改后的完整代码

1. 获取HTML数据部分(无改动)

async function getdata(value){
    let data = "";
    let txtPromise;
    if(value.includes('https')){
        txtPromise = await fetch('https://api.codetabs.com/v1/proxy?quest='+value);
        if (txtPromise.ok) {
           data = await txtPromise.text();
        }
        else{
           data = "not found";
        }
    }
    else
        data = "invalid/null url";
    return data;
}

let urls = ["https://.......","https://......","https://...."];
const responses = []; // 补充声明responses变量
for(var i=0; i<urls.length; i++){ 
    responses.push(getdata(urls[i]));
}

2. 解析处理数据部分(关键修改)

function saveimagetosystem(url){
   return $.ajax({
            url: '/server.php',
            type: 'POST',
            data:{"input":"save image","url":url},
            cache:false,
        })
}

async function processhtmldata(html_data,index){
    if(!html_data.includes("not found") && !html_data.includes("invalid/null url")){
        var parser = new DOMParser();
        var htmldoc = parser.parseFromString(html_data, "text/html");
        var title = htmldoc.querySelector("title").innerText;

        details["title"+" "+index]=title.trim();

        // 修正:判断img元素是否存在
        if(htmldoc.getElementsByTagName("img").length > 0){         
            var url = htmldoc.getElementsByTagName("img")[0].src.trim();
            var response = await saveimagetosystem('https://api.codetabs.com/v1/proxy?quest='+url);
            // 修正拼写:imeg → image
            details["image"+" "+index] = response; 
            // 修正:判断p元素是否存在
            if(htmldoc.getElementsByTagName("p").length > 0){
                for(var l=0; l<htmldoc.getElementsByTagName("p").length-1; l++){                                  
                    details["body"+" "+index+" "+l] =  htmldoc.getElementsByTagName("p")[l].innerText.trim();
                }
            }
        }
    }
}

// 关键修改:用for...of循环保证顺序执行
Promise.all(responses)
  .then(async (htmlfiles) => {
    for (let sl_no = 0; sl_no < htmlfiles.length; sl_no++) {
      await processhtmldata(htmlfiles[sl_no], sl_no);
    }
    // 处理完成后可以在这里访问正确顺序的details对象
    console.log(details);
  })
  .catch(err => console.error('处理失败:', err));
额外优化(并行处理+保序)

如果希望保持并行处理以提高效率,同时保证数据顺序正确,可以收集所有processhtmldata的Promise,再用Promise.all等待全部完成。这种方式下,虽然处理过程是并行的,但每个processhtmldata会把数据写入对应index的字段,最终details的顺序依然和原始URL一致:

Promise.all(responses)
  .then(async (htmlfiles) => {
    const processPromises = htmlfiles.map((file, sl_no) => processhtmldata(file, sl_no));
    await Promise.all(processPromises);
    console.log(details);
  })

内容的提问来源于stack exchange,提问作者user20643052

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 08:31:02