使用Node.js/Cheerio爬取餐厅菜单数据返回空数组求助
Hey there! Let's work through why you're getting an empty array when trying to scrape menu item names from Red Cow's 50th Street menu. From what you shared, it looks like your code cut off mid-way, but let's break down the most common issues and fix this step by step.
Common Causes & Fixes
1. Incorrect CSS Selectors (Most Likely Culprit)
First, you need to make sure your Cheerio selectors match the actual DOM structure of the page. Let's start by verifying what HTML request is fetching:
- Add a line to write the raw HTML to a file so you can inspect it:
fs.writeFileSync("menu.html", html); - Open that file in a browser, then use dev tools to find the exact class/element wrapping each menu item's name and price. Many restaurant menus use containers like
.menu-itemwith child elements for title and price (e.g.,.menu-titleand.menu-price).
2. Incomplete Traversal Logic
Your code might not be properly iterating over menu items. Here's a complete, tested example (adjust selectors based on the actual DOM you inspect):
var request = require("request"); var cheerio = require("cheerio"); var fs = require("fs"); var url = "http://redcowmn.com/50th-street-menu/"; request(url, function(err, response, html) { if (err) { console.error("Oops, failed to fetch the page:", err); return; } // Save raw HTML to debug the page structure fs.writeFileSync("redcow-menu.html", html); var $ = cheerio.load(html); var menuItems = []; // Replace these selectors with the ones you find in the actual DOM $(".menu-section .menu-item").each(function() { var itemName = $(this).find(".menu-item-name").text().trim(); var itemPrice = $(this).find(".menu-item-price").text().trim(); // Only add items that have both a name and price if (itemName && itemPrice) { menuItems.push({ name: itemName, price: itemPrice }); } }); console.log("Scraped Menu Items:", menuItems); fs.writeFileSync("redcow-menu.json", JSON.stringify(menuItems, null, 2)); });
3. Dynamic Content Loading (If Static HTML Lacks the Menu)
If you open redcow-menu.html and don't see any menu items, that means content is loaded dynamically with JavaScript. request only fetches static HTML, so you'll need a tool like Puppeteer to render the page fully before scraping. Here's a quick snippet for that:
const puppeteer = require("puppeteer"); const fs = require("fs"); (async () => { const browser = await puppeteer.launch(); const page = await browser.newPage(); await page.goto("http://redcowmn.com/50th-street-menu/"); const menuItems = await page.evaluate(() => { const items = []; document.querySelectorAll(".menu-section .menu-item").forEach(item => { const name = item.querySelector(".menu-item-name").textContent.trim(); const price = item.querySelector(".menu-item-price").textContent.trim(); if (name && price) items.push({ name, price }); }); return items; }); console.log("Scraped Dynamic Menu:", menuItems); fs.writeFileSync("redcow-dynamic-menu.json", JSON.stringify(menuItems, null, 2)); await browser.close(); })();
Final Quick Tips
- Always inspect the raw HTML first to confirm if content is static or dynamic.
- Double-check your selectors—small typos (like missing a dot for classes) will break everything.
- Add error handling to catch cases where elements don't exist (so you avoid undefined values).
内容的提问来源于stack exchange,提问作者ijones

