You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup解析HTML表格行时,如何跳过含stars类的td.div标签?

解决方案:跳过含stars类的<div>标签解析

问题根源

你之前的错误在于:

  • 使用hasattr(col.div, 'class')判断无效,因为BeautifulSoup的Tag对象本身自带class属性(对应标签的class属性列表),即使标签没有class属性,这个属性也存在(返回None),无法避免KeyError。
  • 直接通过col.div['class']取class属性,当标签没有class时会抛出KeyError。

正确实现代码

替换判断逻辑,使用get('class')方法安全获取class属性(无该属性时返回None),再检查是否包含stars:

def rebuild_row(self, row):
    new_row = []
    for col in row.find_all('td'):
        if col.img:
            continue
        # 精准跳过含stars类的div所在的td
        if col.div:
            div_classes = col.div.get('class')
            if div_classes and 'stars' in div_classes:
                continue
        if col.a:
            new_row.append(self.handle_links(col))
        else:
            if not col.text or not col.text.strip():
                new_row.append(['NaN'])
            else:
                new_text = self.clean_tag_text(col)
                new_row.append(new_text)
    return new_row

逻辑说明

  1. 先判断当前<td>是否包含<div>标签;
  2. 用get('class')获取div的class列表,避免直接取键导致的KeyError;
  3. 若class列表存在且包含stars,则跳过该列的处理;
  4. 其余逻辑保持你原有代码的处理逻辑不变。

内容的提问来源于stack exchange,提问作者Meghan M.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 22:45:32