指定np.float16 dtype后pandas.DataFrame.ffill()报TypeError求解
解决方案:np.float16下ffill()触发TypeError的处理办法
问题根源
np.float16作为半精度浮点数,pandas底层的ffill()函数对其支持存在局限性,部分运算签名未实现,导致直接调用时触发TypeError: No matching signature found,而float32/64是pandas原生支持的浮点类型,因此无此问题。
可行方案
方案1:单列临时转float32填充后转回float16
针对float16类型的列,逐个临时转换为float32执行填充,再转回float16。单列处理内存压力极低(139万行的float32仅约5.5MB),完全符合内存限制:
import pandas as pd import numpy as np for col in df.columns[df.dtypes == np.float16]: df[col] = df[col].astype(np.float32).ffill().astype(np.float16)
方案2:分批次处理ffill()
将数据集拆分为小批次逐段填充,避免全量加载运算触发底层不兼容问题。若存在跨批次的缺失值,可记录上一批次的最后一行值用于填充当前批次开头:
batch_size = 100000 # 根据内存情况调整批次大小 prev_last_row = None for i in range(0, len(df), batch_size): end_idx = min(i + batch_size, len(df)) batch = df.iloc[i:end_idx] if prev_last_row is not None: # 用前一批次最后一行填充当前批次开头的缺失值 batch = batch.combine_first(prev_last_row) batch = batch.ffill().astype(np.float16) df.iloc[i:end_idx] = batch prev_last_row = batch.iloc[-1:]
方案3:读取数据时分块填充(若数据来自文件)
如果数据是从外部文件读取,可在读取阶段分块处理缺失值,避免后续全量操作:
from collections import defaultdict chunk_iter = pd.read_csv('your_data_source.csv', dtype=defaultdict(lambda: np.float16), chunksize=100000) filled_chunks = [] prev_chunk_last = None for chunk in chunk_iter: if prev_chunk_last is not None: chunk = chunk.combine_first(prev_chunk_last) chunk = chunk.ffill() filled_chunks.append(chunk) prev_chunk_last = chunk.iloc[-1:] df = pd.concat(filled_chunks, ignore_index=True)
内容的提问来源于stack exchange,提问作者EvanHong
相关产品推荐
相关产品推荐

