使用Cheerio爬取数据:如何获取<br>标签后的文本内容
Hey there! Let's work through this Cheerio parsing issue you're facing—getting those split odds lines from inside the <a> tag with <br> separators into your newObj properties.
The Core Problem
When you use .text() on that <a> element, Cheerio strips out all HTML tags (including the <br>s) and merges all text content into a single string. That's why your attempts to split on <br> weren't working—those line breaks are already gone by the time you get the text!
Two Practical Solutions
Solution 1: Split the Raw HTML Content
Grab the full HTML inside the <a> tag, split it using <br> as the delimiter, then clean up each segment:
app.get("/nfl", function(req, res) { var results = []; axios.get("http://www.vegasinsider.com/nfl/odds/las-vegas/").then(function(response) { var $ = cheerio.load(response.data); $('span.cellTextHot').each(function(i, element) { var newObj = { time: $(element).text().trim() }; // Extract away and home teams (your existing logic, cleaned up) $(element).parent().children().each(function(i, thing){ if(i === 2){ newObj.awayTeam = $(thing).text().trim(); } else if (i === 4){ newObj.homeTeam = $(thing).text().trim(); } }); // Target the <a> tag containing the odds const oddsAnchor = $(element).parent().next().next().find('a.cellTextNorm'); // Split HTML content by <br> and clean up each segment const oddsSegments = oddsAnchor.html().split('<br>') .map(segment => segment.trim()) // Remove extra spaces/non-breaking spaces .filter(segment => segment !== ''); // Filter out empty strings // Assign segments to your newObj properties if (oddsSegments.length >= 1) newObj.oddsLine1 = oddsSegments[0]; if (oddsSegments.length >= 2) newObj.oddsLine2 = oddsSegments[1]; // Add more lines if you expect additional segments results.push(newObj); // Don't forget to add the object to your results array! }); res.json(results); }).catch(err => { // Add basic error handling for failed requests res.status(500).json({ error: "Failed to fetch or parse data: " + err.message }); }); });
Solution 2: Traverse Text Nodes Directly
Instead of splitting HTML, iterate over the child nodes of the <a> tag and only extract text nodes (ignoring the <br> elements):
// Inside the $('span.cellTextHot').each() loop: const oddsAnchor = $(element).parent().next().next().find('a.cellTextNorm'); const oddsTextNodes = []; oddsAnchor.contents().each((i, node) => { // Only process text nodes (skip elements like <br>) if (node.type === 'text') { const cleanText = $(node).text().trim(); if (cleanText) oddsTextNodes.push(cleanText); } }); // Assign to your newObj properties newObj.oddsLine1 = oddsTextNodes[0] || ''; newObj.oddsLine2 = oddsTextNodes[1] || '';
Quick Tips
- I added
results.push(newObj)because your original code was missing this—without it, yourresultsarray would stay empty! - Adjust the selector
a.cellTextNormif the page's HTML structure changes (always double-check selectors with browser dev tools). - The
.trim()calls clean up extra whitespace and non-breaking spaces ( ) that often show up in scraped content.
内容的提问来源于stack exchange,提问作者kenneth2k1

