You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Cheerio爬取数据:如何获取<br>标签后的文本内容

Fixing Cheerio Parsing to Extract Text After <br> Tags

Hey there! Let's work through this Cheerio parsing issue you're facing—getting those split odds lines from inside the <a> tag with <br> separators into your newObj properties.

The Core Problem

When you use .text() on that <a> element, Cheerio strips out all HTML tags (including the <br>s) and merges all text content into a single string. That's why your attempts to split on <br> weren't working—those line breaks are already gone by the time you get the text!

Two Practical Solutions

Solution 1: Split the Raw HTML Content

Grab the full HTML inside the <a> tag, split it using <br> as the delimiter, then clean up each segment:

app.get("/nfl", function(req, res) {
  var results = [];
  axios.get("http://www.vegasinsider.com/nfl/odds/las-vegas/").then(function(response) {
    var $ = cheerio.load(response.data);
    
    $('span.cellTextHot').each(function(i, element) {
      var newObj = { time: $(element).text().trim() };
      
      // Extract away and home teams (your existing logic, cleaned up)
      $(element).parent().children().each(function(i, thing){
        if(i === 2){
          newObj.awayTeam = $(thing).text().trim();
        } else if (i === 4){
          newObj.homeTeam = $(thing).text().trim();
        }
      });

      // Target the <a> tag containing the odds
      const oddsAnchor = $(element).parent().next().next().find('a.cellTextNorm');
      
      // Split HTML content by <br> and clean up each segment
      const oddsSegments = oddsAnchor.html().split('<br>')
        .map(segment => segment.trim()) // Remove extra spaces/non-breaking spaces
        .filter(segment => segment !== ''); // Filter out empty strings
      
      // Assign segments to your newObj properties
      if (oddsSegments.length >= 1) newObj.oddsLine1 = oddsSegments[0];
      if (oddsSegments.length >= 2) newObj.oddsLine2 = oddsSegments[1];
      // Add more lines if you expect additional segments
      
      results.push(newObj); // Don't forget to add the object to your results array!
    });
    
    res.json(results);
  }).catch(err => {
    // Add basic error handling for failed requests
    res.status(500).json({ error: "Failed to fetch or parse data: " + err.message });
  });
});

Solution 2: Traverse Text Nodes Directly

Instead of splitting HTML, iterate over the child nodes of the <a> tag and only extract text nodes (ignoring the <br> elements):

// Inside the $('span.cellTextHot').each() loop:
const oddsAnchor = $(element).parent().next().next().find('a.cellTextNorm');
const oddsTextNodes = [];

oddsAnchor.contents().each((i, node) => {
  // Only process text nodes (skip elements like <br>)
  if (node.type === 'text') {
    const cleanText = $(node).text().trim();
    if (cleanText) oddsTextNodes.push(cleanText);
  }
});

// Assign to your newObj properties
newObj.oddsLine1 = oddsTextNodes[0] || '';
newObj.oddsLine2 = oddsTextNodes[1] || '';

Quick Tips

  • I added results.push(newObj) because your original code was missing this—without it, your results array would stay empty!
  • Adjust the selector a.cellTextNorm if the page's HTML structure changes (always double-check selectors with browser dev tools).
  • The .trim() calls clean up extra whitespace and non-breaking spaces (&nbsp;) that often show up in scraped content.

内容的提问来源于stack exchange,提问作者kenneth2k1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:16:49