You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas用lambda处理指定列提取HTML内容报错,for循环正常该如何优化?

错误原因
  • 核心原因是df.apply()默认按列(axis=0)遍历,你传入strip_html函数的参数x是整列的Series对象,而非单个单元格的HTML文本。将整列数据转为字符串后BeautifulSoup无法匹配到对应的span.table-content标签,soup.find()返回None,调用get_text()就会触发AttributeError。
  • 次要潜在问题:如果部分单元格的HTML内容缺失目标标签,也会导致相同报错,需要给函数加容错逻辑。
更高效的实现方案

按性能从高到低排序:

方案1:矢量化正则提取(性能最优,比BeautifulSoup方案快5~10倍)

你的提取规则固定,直接用pandas自带的str.extract矢量化操作,不需要逐元素调用解析器,性能最高:

import pandas as pd
import re

target_columns = ["dates", "totals"]
# 预编译正则,匹配目标span内的内容
pattern = re.compile(r"<span class='table-content'>(.*?)</span>")
# 批量处理所有目标列
df[target_columns] = df[target_columns].apply(lambda col: col.str.extract(pattern, expand=False))

方案2:优化BeautifulSoup的列遍历写法(兼容复杂HTML场景)

不要用全表apply,原来的逐列遍历写法本身就比全表apply性能更好,优化函数容错和调用方式即可:

from bs4 import BeautifulSoup

def strip_html(target):
    target = str(target)
    soup = BeautifulSoup(target, "lxml") # 换lxml解析器,比html.parser快30%以上
    span = soup.find("span", class_="table-content")
    return span.get_text() if span else target # 找不到标签返回原值,避免报错

target_columns = ["dates", "totals"]
# 直接传函数名,不需要额外套lambda,减少调用开销
for col in target_columns:
    df[col] = df[col].apply(strip_html)

方案3:修正全表apply的写法(不推荐,性能比逐列处理低)

如果一定要用全表apply,需要指定axis=1逐行处理,同时对每个单元格单独调用函数:

df = df.apply(lambda row: [strip_html(row[col]) if col in target_columns else row[col] for col in df.columns], axis=1, result_type='expand')

内容的提问来源于stack exchange,提问作者SMJune

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 05:48:03