多异步调用下批量爬取多域名商品评分返回未填充数据的解决问询
Hey there! Let's figure out why your getProductDataFromManyDomains is returning empty data and fix it up—this is a super common gotcha with async code!
The Root of the Problem
Chances are your original code looks something like this, right?
function getProductDataFromManyDomains(domains) { const prodData = []; domains.forEach(domain => { scrapWebPage(domain).then(data => { prodData.push(data); }); }); // Oops! This returns before any of the scrapWebPage calls finish return prodData; }
The issue is that scrapWebPage returns a Promise (async work), but your function doesn't wait for those Promises to resolve before returning prodData. The forEach runs synchronously, but the .then() callbacks fire later when the web scraping completes—so you're sending back an empty array before any data is actually added.
Fix 1: Use Promise.all (Parallel Processing, Fastest)
This is the go-to solution for running multiple async tasks at once and waiting for all to finish. It's efficient because all your scraping requests run in parallel.
async function getProductDataFromManyDomains(domains) { // First, create an array of Promises for each domain const scrapePromises = domains.map(domain => { // Wrap in a catch if you want to handle individual failures without breaking everything return scrapWebPage(domain) .catch(error => { console.error(`Failed to scrape ${domain}:`, error); return null; // Return a fallback value so Promise.all doesn't fail entirely }); }); // Wait for ALL Promises to resolve—this gives you the full data array const prodData = await Promise.all(scrapePromises); return prodData; }
Promise.allwaits until every Promise in the array resolves, then returns an array of results in the same order as your input domains.- Adding
.catch()to each individual scrape means if one domain fails, the whole operation doesn't crash—you just get anull(or whatever fallback you want) for that entry.
Fix 2: Serial Processing (For Rate-Limited Sites)
If you're worried about getting blocked for too many parallel requests, you can process domains one at a time with a for...of loop:
async function getProductDataFromManyDomains(domains) { const prodData = []; for (const domain of domains) { try { const data = await scrapWebPage(domain); prodData.push(data); } catch (error) { console.error(`Failed to scrape ${domain}:`, error); prodData.push(null); } // Optional: Add a delay between requests to be nice to the server // await new Promise(resolve => setTimeout(resolve, 1000)); } return prodData; }
This is slower, but safer if the target sites have strict rate limits. Each scrape finishes before the next one starts.
Key Takeaway
Any time you're working with async functions (Promises), you need to explicitly wait for them to resolve before accessing their results. Using async/await with Promise.all is the cleanest pattern for bulk async operations like this.
内容的提问来源于stack exchange,提问作者stackustack

