导出清洗后数据集至CSV时触发AttributeError错误求助
问题分析
错误 AttributeError: 'tuple' object has no attribute 'to_csv' 的根源很明确:
你最初将data_cleaned_dropped定义为Pandas DataFrame对象(通过data_cleaned.drop(columns=columns_to_drop)生成),但后续代码里错误地把它重新赋值为该DataFrame的shape属性——而shape返回的是一个元组(格式为(行数, 列数)),元组不具备to_csv方法,因此触发报错。
问题出在这段代码:
# 此处将原DataFrame变量覆盖为shape元组 data_cleaned_dropped = data_cleaned_dropped.shape
修复方案
只需避免覆盖原DataFrame变量,将获取维度的操作转移到新变量中即可。修正后的完整代码片段如下:
# Columns to drop based on the missing percentage columns_to_drop = [ 'Nazwa', 'Czas ostatniej modyfikacji', 'Jeżeli pominięto branżę, na której się znasz dopisz ją:', 'Jeżeli jest obszar, na którym się znasz i chcesz go wykorzystać, dopisz go:', 'Masz jakiś pomysł na projekt data? Napisz nam o tym. Jeżeli to nie ten moment, pozostaw puste pole.' ] # Dropping the specified columns from the cleaned dataset data_cleaned_dropped = data_cleaned.drop(columns=columns_to_drop) # Checking the size of the cleaned dataset after dropping columns dataset_size = data_cleaned_dropped.shape print("Number of rows:", dataset_size[0]) print("Number of columns:", dataset_size[1]) # Display the first few rows of the updated dataset to confirm the columns are dropped print(data_cleaned_dropped.head()) # Checking the size of the original dataset original_dataset_size = data.shape print("Number of rows:", original_dataset_size[0]) print("Number of columns:", original_dataset_size[1]) # 关键修改:用新变量存储维度信息,不覆盖原DataFrame new_dataset_size = data_cleaned_dropped.shape print("New number of rows:", new_dataset_size[0]) print("New number of columns:", new_dataset_size[1]) # Save the cleaned and updated dataset to a CSV file output_file_path = r'D:\Users\Desktop\ProjektAnalitiks\data\02_Interim\dropped_rows_columns.csv' data_cleaned_dropped.to_csv(output_file_path, index=False) print(f"Dataset saved to {output_file_path}")
核心提示
- 不要随意覆盖存储核心数据集的变量,避免丢失DataFrame对象
- Pandas中
DataFrame.shape仅用于获取数据集的维度信息,返回的元组无法执行任何数据导出、清洗等操作
内容的提问来源于stack exchange,提问作者rejent
相关产品推荐
相关产品推荐

