如何解决sum()函数将数值当作字符串拼接而非求和的问题
问题:sum()函数将数值当作字符串拼接而非求和
原代码
df = my_dict[key] print(df) inst_row = df.iloc[3] print(inst_row) inst_sum = inst_row.sum() print(inst_sum)
运行输出
df 0 1 2 ... 1366 1367 1368 ID 185 278 451 ... 199622 199746 199940 A 484406.166667 493122.5 481308.166667 ... 481924.5 484088.333333 482801.166667 B 4,096 4,096 4,096 ... 4,096 4,096 4,096 C 87867392.0 87867392.0 87867392.0 ... 87867392.0 87867392.0 87867392.0 inst_row [4 rows x 1369 columns] 0 87867392.0 1 87867392.0 2 87867392.0 3 87867392.0 4 87867392.0 ... 1364 87867392.0 1365 87867392.0 1366 87867392.0 1367 87867392.0 1368 87867392.0 Name: C, Length: 1369, dtype: object inst_sum 87867392.087867392.087867392.0......87867392.087867392.0
原因分析
从输出可见inst_row的 dtype 是object,说明这一行的元素本质是字符串类型。Pandas 对字符串执行sum()时,默认会做字符串拼接操作,而非数值求和。
解决方案
求和前先将该行数据转换为数值类型,以下是两种常用实现方式:
方法1:用pd.to_numeric()转换(容错性高,推荐)
import pandas as pd inst_row = df.iloc[3] # errors='coerce'会把无法转换的值设为NaN,避免转换失败报错 inst_row_numeric = pd.to_numeric(inst_row, errors='coerce') inst_sum = inst_row_numeric.sum() print(inst_sum)
方法2:用astype()强制转换(适合确认所有值都能转成数值的场景)
inst_row = df.iloc[3] inst_sum = inst_row.astype(float).sum() print(inst_sum)
特殊情况处理:数据带千分位逗号(比如B列的4,096)
需要先去掉逗号再转换:
b_row = df.iloc[2].str.replace(',', '').astype(float) b_sum = b_row.sum() print(b_sum)
内容的提问来源于stack exchange,提问作者mahmood
相关产品推荐
相关产品推荐

