You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas数据框归一化时出现str与float减法类型错误

问题背景

neg_ctl_df数据框存储阴性对照数据,coding_gene_df数据框存储目标基因数据。需要对每个样本执行归一化操作——减去该样本内阴性对照的中位数。samples和neg_ctl_median均为<class 'pandas.core.series.Series'>类型。

执行代码
import pandas as pd

# 阴性对照归一化:减去患者样本内阴性对照的中位数
neg_ctl_median = neg_ctl_df.iloc[:,-29:].median()

for gene, samples in coding_gene_df.iloc[:,-29:].iterrows():
  norm_val = samples - neg_ctl_median
  print(norm_val)
报错信息
---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
/usr/local/lib/python3.7/dist-packages/pandas/core/ops/array_ops.py in _na_arithmetic_op(left, right, op, is_cmp)
    165     try:
--> 166         result = func(left, right)
    167     except TypeError:

9 frames
TypeError: unsupported operand type(s) for -: 'str' and 'float'

During handling of the above exception, another exception occurred:

TypeError                                 Traceback (most recent call last)
/usr/local/lib/python3.7/dist-packages/pandas/core/ops/array_ops.py in _masked_arith_op(x, y, op)
    110         # See GH#5284, GH#5035, GH#19448 for historical reference
    111         if mask.any():
--> 112             result[mask] = op(xrav[mask], yrav[mask])
    113 
    114     else:

TypeError: unsupported operand type(s) for -: 'str' and 'float'
数据示例

样本数据(coding_gene_df.iloc[1:10,-29:-27].to_dict()):

{'12h_P1_T4_TimeC2_PIDC4_Non-Survivor': {'CNTN2': '6.35',
  'KCNA2': '5.29',
  'LOC79160': '5.99',
  'PTGIS': '5.66',
  'TTTY11': '3.91',
  'VPS4B': '9.68',
  'XRCC1': '9.09',
  'ZC3HC1': '7.19',
  'ZFAS1': '8.68'},
 '48h_P1_T6_TimeC3_PIDC1_Non-Survivor': {'CNTN2': '6.6',
  'KCNA2': '5.36',
  'LOC79160': '6.18',
  'PTGIS': '5.54',
  'TTTY11': '3.92',
  'VPS4B': '9.51',
  'XRCC1': '9.15',
  'ZC3HC1': '7.05',
  'ZFAS1': '8.46'}}

阴性对照数据(neg_ctl_df.iloc[1:10,-29:-27].to_dict()):

{'12h_P1_T4_TimeC2_PIDC4_Non-Survivor': {'---': '8.45'},
 '48h_P1_T6_TimeC3_PIDC1_Non-Survivor': {'---': '8.16'}}
数据类型
print(type(neg_ctl_median))
<class 'pandas.core.series.Series'>

print(type(samples))
<class 'pandas.core.series.Series'>
问题原因与解决方案

问题根源

从数据示例可见,coding_gene_df和neg_ctl_df中的数值均为字符串类型(如'6.35'、'8.45'),而中位数计算结果为浮点数,字符串与浮点数无法直接执行减法运算,这就是报错的核心原因。

修复步骤

  1. 转换数据类型:将两个数据框中需要计算的列转换为数值类型(float),用pd.to_numeric处理时可将无效值转为NaN。
  2. 优化计算逻辑:无需循环遍历行,直接对整个数据框做向量运算,效率更高。

修复后的代码

import pandas as pd

# 处理阴性对照数据:转换为数值类型,无效值转为NaN
neg_ctl_processed = neg_ctl_df.iloc[:, -29:].apply(pd.to_numeric, errors='coerce')
# 计算每个样本的阴性对照中位数
neg_ctl_median = neg_ctl_processed.median()

# 处理目标基因数据:转换为数值类型
coding_gene_processed = coding_gene_df.iloc[:, -29:].apply(pd.to_numeric, errors='coerce')
# 执行归一化:每个样本的基因值减去对应阴性对照中位数
norm_df = coding_gene_processed - neg_ctl_median

print(norm_df)

说明

  • pd.to_numeric(errors='coerce')会把无法转换为数值的内容转为NaN,避免转换报错。
  • 直接对数据框执行减法运算,pandas会自动按列(样本)对齐,比循环遍历效率高得多,尤其适合大数据量场景。

内容的提问来源于stack exchange,提问作者melolilili

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 00:54:33