You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用JavaScript过滤HTML爬取数据并提取指定class元素内容

解决方法

你可以直接通过cheerio的类选择器定位所有fight-date元素,再遍历提取文本即可,不需要先将上层节点转为字符串处理,修改后的代码如下:

const scheduleBoxingScene = "https://www.boxingscene.com/schedule";
axios(scheduleBoxingScene).then((response) => {
  const html = response.data;
  const $ = cheerio.load(html);
  const fightDateList = [];
  // 直接定位所有fight-date元素并遍历
  $(".body-container .schedule .fight-date").each((index, element) => {
    // 提取文本并清除首尾多余空白
    const dateText = $(element).text().trim();
    fightDateList.push(dateText);
  })
  // 输出数组结果,也可直接转为JSON格式
  console.log(fightDateList);
  console.log(JSON.stringify(fightDateList, null, 2))
})
  • 直接使用.body-container .schedule .fight-date选择器可以一步定位到所有目标日期节点,无需提前提取上层div内容
  • 调用.each()方法遍历所有匹配节点,逐个获取元素文本后用.trim()去除首尾空白字符,即可过滤掉冗余的换行、空格
  • 最终得到的日期数组可直接用JSON.stringify()转为标准JSON格式存储

内容的提问来源于stack exchange,提问作者Neto Silveira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 22:15:02