使用Pandas修改CSV内容失败:DataFrame修改未保存至新CSV文件
问题分析与解决
你的代码存在三个关键问题,导致修改无法同步到CSV文件:
1. 未真正修改DataFrame内容
你通过row = housing.at[index, headers[header_index]]把DataFrame中的值取出赋值给临时变量row,后续修改row = ''只是改变了这个变量的内容,并没有同步回原DataFrame。必须直接对housing.at[...]赋值才能修改DataFrame。
2. 类型判断逻辑错误
if row is not int or row is not float这个条件永远为真——任何值都不可能同时是int和float,不管row是什么类型,条件都会触发。正确的判断应该用isinstance()函数检查是否属于数值类型:
if not isinstance(row, (int, float))
3. 嵌套循环效率低下
Pandas设计初衷是矢量化操作,嵌套循环遍历每一行每一列的方式在数据量大时会非常慢,不推荐使用。
修正后的代码(循环方式)
import pandas as pd housing = pd.read_csv(csv_path) headers = ['Sold Price', 'Longitude', 'Latitude', 'Land Size', 'Total Bedrooms', 'Total Bathrooms', 'Parking Spaces'] for header in headers: for index in housing.index: value = housing.at[index, header] if not isinstance(value, (int, float)): # 直接修改DataFrame中的对应位置 housing.at[index, header] = '' housing.to_csv('propertyupdated.csv')
更高效的矢量化实现(推荐)
用applymap批量处理指定列,避免嵌套循环:
import pandas as pd housing = pd.read_csv(csv_path) headers = ['Sold Price', 'Longitude', 'Latitude', 'Land Size', 'Total Bedrooms', 'Total Bathrooms', 'Parking Spaces'] # 批量将非数值类型替换为空字符串 housing[headers] = housing[headers].applymap(lambda x: '' if not isinstance(x, (int, float)) else x) housing.to_csv('propertyupdated.csv')
或者用pd.to_numeric处理,适合将非数值转为空的场景:
import pandas as pd housing = pd.read_csv(csv_path) headers = ['Sold Price', 'Longitude', 'Latitude', 'Land Size', 'Total Bedrooms', 'Total Bathrooms', 'Parking Spaces'] for header in headers: # 先将非数值转为NaN,再替换为空字符串 housing[header] = pd.to_numeric(housing[header], errors='coerce').fillna('') housing.to_csv('propertyupdated.csv')
内容的提问来源于stack exchange,提问作者CountDOOKU
相关产品推荐
相关产品推荐

