You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas按指定条件根据另一列填充Date列空值?

Pandas填充Date列空值的类型错误解决方法

问题场景

现有如下DataFrame,需求为:当Area列值为West或North时,用Year列值拼接'0601'填充Date列的空单元格。尝试两段代码均触发相同类型错误,代码及报错如下:

测试代码

from io import StringIO
import pandas as pd
import numpy as np  # 原代码遗漏numpy导入

csvfile = StringIO(
"""
Name    Area    Date    Year
David   West        2014
Mike    North   20220919    2022
Kate    West        2017
Lilly   East    20221226    2022
Peter   North   20221226    2022
Cara    Middle      2016

""")

df = pd.read_csv(csvfile, sep='\t', engine='python')

L1 = ['West','North']
m1 = df['Date'].isnull()
m2 = df['Area'].isin(L1)

df['Date'] = df['Date'].mask(m1 & m2, df['Year'] + '0601')      # Try_1

df['Date'] = np.where(np.where(m1 & m2, df['Year'] + '0601'))   # Try_2

报错信息(中文翻译版)

回溯(最近的调用最后):
  文件 "C:\Python38\lib\site-packages\pandas\core\ops\array_ops.py", 第142行, 在 _na_arithmetic_op中
    result = expressions.evaluate(op, left, right)
  文件 "C:\Python38\lib\site-packages\pandas\core\computation\expressions.py", 第235行, 在 evaluate中
    return _evaluate(op, op_str, a, b)  # type: ignore[misc]
  文件 "C:\Python38\lib\site-packages\pandas\core\computation\expressions.py", 第69行, 在 _evaluate_standard中
    return op(a, b)
numpy.core._exceptions.UFuncTypeError: 通用函数(ufunc) 'add' 没有匹配签名的循环:(dtype('<U21'), dtype('<U21')) -> dtype('<U21')

处理上述异常期间,又触发新的异常:

回溯(最近的调用最后):
  文件 "C:\My Documents\Scripts\(Desktop) WSS 20200323\GG.py", 第336行, 在 <module>中
    df['Date'] = np.where(np.where(m1 & m2, df['Year'] + '0601'))                   # try 2
  文件 "C:\Python38\lib\site-packages\pandas\core\ops\common.py", 第65行, 在 new_method中
    return method(self, other)
  文件 "C:\Python38\lib\site-packages\pandas\core\arraylike.py", 第89行, 在 __add__中
    return self._arith_method(other, operator.add)
  文件 "C:\Python38\lib\site-packages\pandas\core\series.py", 第4998行, 在 _arith_method中
    result = ops.arithmetic_op(lvalues, rvalues, op)
  文件 "C:\Python38\lib\site-packages\pandas\core\ops\array_ops.py", 第189行, 在 arithmetic_op中
    res_values = _na_arithmetic_op(lvalues, rvalues, op)
  文件 "C:\Python38\lib\site-packages\pandas\core\ops\array_ops.py", 第149行, 在 _na_arithmetic_op中
    result = _masked_arith_op(left, right, op)
  文件 "C:\Python38\lib\site-packages\pandas\core\ops\array_ops.py", 第111行, 在 _masked_arith_op中
    result[mask] = op(xrav[mask], y)
numpy.core._exceptions.UFuncTypeError: 通用函数(ufunc) 'add' 没有匹配签名的循环:(dtype('<U21'), dtype('<U21')) -> dtype('<U21')

错误原因

  1. 类型不匹配:Year列是整数类型,无法直接与字符串'0601'用+拼接,Python不支持数值与字符串的加法运算,触发类型错误。
  2. Try_2的np.where用法错误:np.where需要三个参数:条件表达式、条件满足时的值、条件不满足时的值,原代码只传了嵌套的np.where,参数缺失。

正确写法

方法1:修正mask方法

将Year列转为字符串后再拼接,代码如下:

df['Date'] = df['Date'].mask(m1 & m2, df['Year'].astype(str) + '0601')

方法2:修正np.where方法

补全np.where的参数,同时转换Year类型:

df['Date'] = np.where(m1 & m2, df['Year'].astype(str) + '0601', df['Date'])

完整可运行代码

from io import StringIO
import pandas as pd
import numpy as np

csvfile = StringIO(
"""
Name    Area    Date    Year
David   West        2014
Mike    North   20220919    2022
Kate    West        2017
Lilly   East    20221226    2022
Peter   North   20221226    2022
Cara    Middle      2016

""")

df = pd.read_csv(csvfile, sep='\t', engine='python')

L1 = ['West','North']
m1 = df['Date'].isnull()
m2 = df['Area'].isin(L1)

# 二选一即可
# 方法1
df['Date'] = df['Date'].mask(m1 & m2, df['Year'].astype(str) + '0601')

# 方法2,若用这个则注释掉方法1
# df['Date'] = np.where(m1 & m2, df['Year'].astype(str) + '0601', df['Date'])

print(df)

运行后输出结果:

Name    Area      Date  Year
0  David    West  20140601  2014
1   Mike   North  20220919  2022
2   Kate    West  20170601  2017
3  Lilly    East  20221226  2022
4  Peter   North  20221226  2022
5   Cara  Middle       NaN  2016

内容的提问来源于stack exchange,提问作者Mark K

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 11:15:36