使用Puppeteer的page.evaluate时遇TypeError错误求助
Hey there, let's break down why you're hitting that TypeError and get your EWRC results scraper working properly.
What's Causing the Error?
The crash happens because your grabFromRow function tries to access innerText on a null value. That means when you run row.querySelector(td.${classname}), it can't find a matching element for one of your target classes—so it returns null, and calling innerText on that crashes the script.
Let's walk through the specific issues and fixes:
1. Typo in Row Selector
Your row selector tr.table_sude has a typo! Looking at the EWRC page, the driver table rows use the class table_side (not sude). Using the wrong class means you're either selecting no rows at all, or incorrect rows that don't have the TDs you're trying to scrape.
2. Wrong Selector for Driver Names
When you try to grab the driver name with grabFromRow(tr, 'a'), you're telling the function to look for a <td class="a">—which doesn't exist. The driver name lives inside a <td class="points-name"> element, specifically in an <a> tag nested inside that TD.
3. Insufficient Page Load Waiting
Using waitUntil: 'domcontentloaded' only waits for the initial DOM to load, but the page might still be rendering dynamic content. If your script runs before the table rows are fully loaded, it'll try to query elements that don't exist yet.
Fixed Code
Here's the updated scraper with all the fixes, plus extra safeguards to prevent future crashes:
const puppeteer = require('puppeteer'); async function getChampTable(year) { try { const browser = await puppeteer.launch({ headless: true }); // Set to false to see the browser for debugging const page = await browser.newPage(); const url = `https://www.ewrc-results.com/season/${year}/1-wrc/`; // Wait for network activity to settle to ensure full page load await page.goto(url, { waitUntil: 'networkidle2' }); // Wait explicitly for the driver rows to appear before scraping await page.waitForSelector('tr.table_side'); const driverTable = await page.evaluate(() => { // Add error handling to avoid crashes if elements are missing const grabFromRow = (row, selector) => { const element = row.querySelector(selector); return element ? element.innerText.trim() : ''; }; const DRIVER_ROW_SELECTOR = 'tr.table_side'; // Fixed typo const data = []; const driverRows = document.querySelectorAll(DRIVER_ROW_SELECTOR); for (const tr of driverRows) { data.push({ position: grabFromRow(tr, 'td.points-pos'), name: grabFromRow(tr, 'td.points-name a'), // Correct selector for driver name pointsTotal: grabFromRow(tr, 'td.points-total') }); } // Filter out any empty entries to clean up results return data.filter(item => item.position && item.name); }); console.log(JSON.stringify(driverTable, null, 2)); await browser.close(); // Always close the browser to avoid leftover processes } catch (error) { console.error('Scraping failed:', error); } } getChampTable(1995);
Key Improvements:
- Fixed row selector typo:
table_sideinstead oftable_sude - Robust
grabFromRow: Now checks if the element exists before accessinginnerText, returning an empty string instead of crashing - Correct name selector: Targets the actual
<a>tag inside thepoints-nameTD - Better load waiting: Uses
networkidle2to wait for most network activity to stop, pluswaitForSelectorto ensure the rows are rendered - Cleanup: Added
browser.close()to prevent hanging browser processes, and filtered out empty entries
Debugging Tip
If you run into issues again, set headless: false in puppeteer.launch()—this will open a visible browser window so you can inspect the page and verify your selectors match the actual elements.
内容的提问来源于stack exchange,提问作者James Cowell

