升级Pandas 2后设置float16类型索引报错问题排查
Pandas 2中float16索引报错问题解决
问题背景
升级至Pandas 2版本后,执行以下代码时出现报错。此前为降低内存占用使用float16类型一直正常,升级后无法运行:
执行代码
file='test_read_float16.csv' df=pd.read_csv(file,sep='\t') df # 输出数据: # Depth 2023-05-12 2023-05-12 0:20 2023-05-12 0:36 # 0 0 19.593750 19.296875 20.59375 # 1 1 23.296875 21.906250 21.00000 # 2 2 112.187500 112.187500 111.68750 # 3 3 180.750000 180.750000 180.25000 # 4 4 187.375000 188.500000 188.12500
df=df.astype('float16',errors='ignore') df=df.set_index('Depth') df
报错信息
NotImplementedError Traceback (most recent call last) cnrl\users\yongnual\Data\Spyder_workplace\DTS_dashboard\DTS_dashboard_v201_injectorDateRange_reducememory_seeqcrossplot_calcfluidlevel_crossplotseeq_crossplotspm_b12dtscrossplot_pandas2_parquet.ipynb Cell 205 line 2 1 df=df.astype('float16',errors='ignore') ----> 2 df=df.set_index('Depth') 3 df File c:\Anaconda\envs\dash2\lib\site-packages\pandas\core\frame.py:5915, in DataFrame.set_index(self, keys, drop, append, inplace, verify_integrity) 5907 if len(arrays[-1]) != len(self): 5908 # check newest element against length of calling frame, since 5909 # ensure_index_from_sequences would not raise for append=False. 5910 raise ValueError( 5911 f"Length mismatch: Expected {len(self)} rows, " 5912 f"received array of length {len(arrays[-1])}" 5913 ) -> 5915 index = ensure_index_from_sequences(arrays, names) 5917 if verify_integrity and not index.is_unique: 5918 duplicates = index[index.duplicated()].unique() File c:\Anaconda\envs\dash2\lib\site-packages\pandas\core\indexes\base.py:7067, in ensure_index_from_sequences(sequences, names) 7065 if names is not None: 7066 names = names[0] -> 7067 return Index(sequences[0], name=names) 7068 else: 7069 return MultiIndex.from_arrays(sequences, names=names) ... 578 # asarray_tuplesafe does not always copy underlying data, 579 # so need to make sure that this happens 580 data = data.copy() NotImplementedError: float16 indexes are not supported
问题原因
Pandas 2并没有取消对float16的支持,而是不允许将float16类型的数据设置为索引。报错信息已明确提示float16 indexes are not supported。旧版本Pandas可能未严格校验索引的浮点类型精度问题,因此未触发报错;而Pandas 2收紧了索引类型限制,因为float16的精度不足,作为索引会引发潜在的匹配错误、索引失效等问题。
解决方案
方案1:先设置索引,再转换其他列为float16
优先保持索引列的原有精度(如整数或float64),仅将数据列转换为float16以节省内存:
file='test_read_float16.csv' df=pd.read_csv(file,sep='\t') # 先将Depth设为索引,再转换其他数据列 df = df.set_index('Depth') df = df.astype('float16', errors='ignore')
方案2:转换索引列为更高精度浮点型后再设置索引
如果必须保留Depth列的浮点属性,可将其转换为float32或float64后再设为索引,其他列仍用float16:
file='test_read_float16.csv' df=pd.read_csv(file,sep='\t') # 将Depth转换为float32(精度足够且内存占用远低于float64) df['Depth'] = df['Depth'].astype('float32') df = df.set_index('Depth') # 其他数据列转换为float16 df = df.astype('float16', errors='ignore')
内容的提问来源于stack exchange,提问作者roudan
相关产品推荐
相关产品推荐

