You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup提取标签class属性报KeyError如何解决

问题说明

尝试从td标签提取class属性值时触发KeyError: 'class'报错,原实现代码如下:

def getLine (self):
    dat = BeautifulSoup(self.data, "html.parser")
    tags = dat.find_all("tr")
    for current in tags:
        line = current.findChildren("td", recursive=False)
        for currentLine in line:
            # print (currentLine)
            clase = currentLine["class"] # <-- 报错行
            if clase is not None and "result" in clase:
                valor = Line()
                valor.name = line.text

报错栈信息:

File "C:\DataProgb\urlwrapper.py", line 43, in getLine
    clase = currentLine["class"] #Trying to extract attribute class
\AppData\Local\Programs\Python\Python39\lib\site-packages\bs4\element.py", line 1519, in __getitem__
    return self.attrs[key]
KeyError: 'class'

目标逻辑为:判断标签是否存在class属性,且属性值包含"result"时执行后续业务操作。

报错原因

BeautifulSoup的标签对象使用标签["属性名"]语法取值时,逻辑和Python字典按键取值完全一致:标签不存在对应属性时会直接抛出KeyError,而非返回None。遍历过程中部分td标签未定义class属性,直接用方括号取值就会触发该报错。

另外代码存在一处隐藏逻辑错误:line是findChildren返回的td标签列表,列表对象不存在.text属性,即便class取值不报错,valor.name = line.text行也会抛出异常,此处应该取当前遍历到的单个td对象currentLine的文本内容。

解决方法

方法1:使用.get()方法安全读取属性

和字典的.get()用法一致,属性不存在时默认返回None,不会触发报错,修正后代码如下:

def getLine (self):
    dat = BeautifulSoup(self.data, "html.parser")
    tags = dat.find_all("tr")
    for current in tags:
        tds = current.findChildren("td", recursive=False)
        for current_td in tds:
            clase = current_td.get("class")
            # BeautifulSoup会将class这类多值属性自动解析为列表,直接判断成员即可,无需单独判空
            if clase and "result" in clase:
                valor = Line()
                valor.name = current_td.text # 修正为取当前单个td的文本

补充说明:HTML的class属性支持同时配置多个值(比如class="result highlight"),BeautifulSoup会自动将这类多值属性拆分为列表返回,直接用"result" in clase即可完成匹配,无需手动拼接字符串。

方法2:查找标签时直接过滤class条件

无需手动遍历判断属性,可以直接在查找标签时传入class过滤规则,代码更简洁,出错概率更低:

def getLine (self):
    dat = BeautifulSoup(self.data, "html.parser")
    tags = dat.find_all("tr")
    for current in tags:
        # 直接筛选当前tr下、class包含result的直接子td
        for current_td in current.findChildren("td", recursive=False, class_="result"):
            valor = Line()
            valor.name = current_td.text
            # 后续业务逻辑直接写在这里

如果习惯用CSS选择器,也可以直接用select方法一步筛选目标标签:

def getLine (self):
    dat = BeautifulSoup(self.data, "html.parser")
    # 匹配所有tr的直接子td中,class包含result的标签
    for current_td in dat.select("tr > td.result"):
        valor = Line()
        valor.name = current_td.text

内容的提问来源于stack exchange,提问作者Joseph

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 00:01:08