如何在Python循环前一次性获取用户输入指定pandas写入csv的模式
解决方案
核心思路是在循环外新增两个全局标记变量,记录用户的选择和是否已经触发过写入逻辑,避免重复弹窗,具体修改如下:
核心修改点
- 循环开始前初始化
user_write_choice存储用户写入选择,is_first_write标记是否是第一次触发写入操作 - 第一次触发写入逻辑时,若目标文件非空,仅询问一次用户选择,后续直接复用该选择
- 若用户选择覆盖模式,第一次写入使用
w模式覆盖原文件,后续所有写入自动切换为a追加模式,避免后续写入覆盖之前的新数据
修改后完整代码
import os import pandas as pd # 初始化全局标记,放在循环外面 user_write_choice = None is_first_write = True i = 0 for index, row in read_file.iterrows(): # 拆分Case只做一次,减少重复计算 case_parts = row['Case'].split('-') first, second, third, fourth, fifth = case_parts if first == 'X01' and second == '01' and fourth == '04': i += 1 Ax = float(row['Ax']) Ay = float(row['Ay']) Az = float(row['Az']) ENT = float(row['ENT']) Ips = (Ax**2 + Ay**2 + Az**2)**0.5 beta = float(row['beta']) date = row['Date'].replace("/", "-") totalP = float(row['Total.P']) data = pd.DataFrame({'Case': [str(row['Case'])], 'ENT': [ENT], 'total P': [totalP]}, index = [i]) filename = 'curve_{}-{}-{}-{}_I{}-B{}-D{}.csv'.format(first,second,third,fourth,round(Ips, 2),beta,date) # 先判断文件是否存在,避免os.path.getsize报错 file_exists = os.path.exists(filename) and os.path.getsize(filename) > 0 # 仅第一次写入且文件非空时询问用户 if file_exists and is_first_write: print('Curve file not empty.') while True: inp = input('Do you want to: A) Append the file. B) Overwrite the file. [A/B]? : ') if inp in ['A', 'B']: user_write_choice = inp break # 确定当前写入模式 if user_write_choice == 'B' and is_first_write: # 覆盖模式下第一次写入用w,之后自动切追加 mode = 'w' print('Overwriting... ') else: mode = 'a' if is_first_write and not file_exists: print('Creating new file...') else: print('Appending... ') # 写入文件,注意如果是新文件第一次写要保留表头,后续追加不要表头 data.to_csv(filename, mode=mode, header=is_first_write, index_label='index') # 第一次写入完成后修改标记 if is_first_write: is_first_write = False
额外优化建议
你当前逐行遍历的方式在数据量大的时候性能较差,可以直接用pandas矢量化操作一次性过滤所有符合条件的行,处理完成后一次性写入文件,代码更简洁效率也更高:
# 拆分Case列为多列 read_file[['first', 'second', 'third', 'fourth', 'fifth']] = read_file['Case'].str.split('-', expand=True) # 过滤符合条件的行 filter_df = read_file[(read_file['first'] == 'X01') & (read_file['second'] == '01') & (read_file['fourth'] == '04')].copy() # 计算需要的字段 filter_df['Ips'] = (filter_df['Ax'].astype(float)**2 + filter_df['Ay'].astype(float)**2 + filter_df['Az'].astype(float)**2)**0.5 filter_df['date'] = filter_df['Date'].str.replace('/', '-') # 生成文件名(因为符合条件的行参数一致,取第一行即可) first_row = filter_df.iloc[0] filename = 'curve_{}-{}-{}-{}_I{}-B{}-D{}.csv'.format( first_row['first'], first_row['second'], first_row['third'], first_row['fourth'], round(first_row['Ips'], 2), float(first_row['beta']), first_row['date'] ) # 生成输出格式的DataFrame output_df = pd.DataFrame({ 'Case': filter_df['Case'], 'ENT': filter_df['ENT'].astype(float), 'total P': filter_df['Total.P'].astype(float) }).reset_index(drop=True) output_df.index += 1 # 序号从1开始 # 写入前询问用户 file_exists = os.path.exists(filename) and os.path.getsize(filename) > 0 if file_exists: print('Curve file not empty.') while True: inp = input('Do you want to: A) Append the file. B) Overwrite the file. [A/B]? : ') if inp in ['A', 'B']: break if inp == 'A': print('Appending... ') output_df.to_csv(filename, mode='a', header=False, index_label='index') else: print('Overwriting... ') output_df.to_csv(filename, mode='w', index_label='index') else: print('Creating new file...') output_df.to_csv(filename, mode='w', index_label='index')
内容的提问来源于stack exchange,提问作者Bean_from_accounts
相关产品推荐
相关产品推荐

