如何在Puppeteer中实现Selenium getText()的等效功能?
Great question—this whitespace normalization quirk is a common pain point when moving from Selenium to Puppeteer. The good news is you can replicate Selenium's getText() behavior with a combination of innerText and some simple string processing.
Why the Discrepancy?
Selenium's getText() returns visible, normalized text: it accounts for how elements are rendered (e.g., adding spaces between block-level elements) and collapses any sequence of whitespace (spaces, tabs, newlines) into a single space, plus trims leading/trailing whitespace.
Puppeteer's textContent, on the other hand, just concatenates all text nodes directly from the DOM, ignoring rendering rules and whitespace normalization—hence why you get Unit Tests15 instead of Unit Tests 15.
Solution: Replicate Selenium's Normalization
The key is to use innerText (which reflects rendered text, including line breaks between block elements) and then normalize the whitespace to match Selenium's output.
1. Helper Function for Reusability
Since you're migrating thousands of lines, create a reusable helper function to avoid repeating code:
async function getText(elementHandle) { // Get rendered text (matches what a user would see) const rawText = await elementHandle.evaluate(el => el.innerText); // Normalize whitespace: collapse multiple spaces/newlines into one, trim edges return rawText.replace(/\s+/g, ' ').trim(); }
2. Usage Example
For your specific HTML structure:
<div test-id="something"> <div>Unit Tests</div> <div>15</div> </div>
You'd use it like this:
// Get the element handle const element = await page.$('[test-id="something"]'); // Get normalized text (matches Selenium's getText()) const normalizedText = await getText(element); // Now this comparison will pass! console.log(normalizedText === "Unit Tests 15"); // true
3. Inline Version (if you prefer)
If you want to avoid a helper function for one-off cases:
const normalizedText = await page.$eval('[test-id="something"]', el => el.innerText.replace(/\s+/g, ' ').trim() );
Key Notes
- Visibility:
innerTextignores hidden elements (just like Selenium'sgetText()), whereastextContentincludes hidden text. This ensures you're only getting visible text as intended. - Whitespace Handling: The regex
/\s+/gmatches any sequence of whitespace characters (spaces, tabs, newlines) and replaces them with a single space. Trimming removes any leading/trailing whitespace that might be present.
This approach should work for almost all cases where you need to match Selenium's getText() output in Puppeteer.
内容的提问来源于stack exchange,提问作者Austin France

