CSV邮箱哈希脚本报错:list index out of range问题求助
排查SHA256哈希CSV邮箱时的list index out of range错误
错误根源
- 空行或无效行:CSV文件里存在空行(比如末尾的空白行)时,
customer_email会变成空列表,访问customer_email[0]自然触发索引越界。 - 行数据异常:部分行没有第一列数据(比如只有分隔符但无内容,或者列结构和表头不匹配),导致列表长度为0。
- 冗余文件操作:
with语句会自动关闭文件,原代码里的f_input.close()和f_output.close(还少了括号)完全多余,属于不必要的语法问题。
修复后的代码
import hashlib import csv import glob def hash(text): return hashlib.sha256(text.encode('UTF-8')).hexdigest() def hash_file(input_file_name, output_file_name): with open(input_file_name, newline='') as f_input, open(output_file_name, 'w', newline='') as f_output: csv_input = csv.reader(f_input) csv_output = csv.writer(f_output) # 处理空文件情况 try: header = next(csv_input) csv_output.writerow(header) except StopIteration: print(f"文件 {input_file_name} 是空的,跳过") return count = 0 print(count) for customer_email in csv_input: # 跳过空行或无第一列的行 if not customer_email or len(customer_email) < 1: print(f"跳过无效行: {customer_email}") continue # 清理邮箱内容,处理空字符串 email = customer_email[0].strip() if not email: print(f"跳过空邮箱") csv_output.writerow([""]) continue csv_output.writerow([hash(email)]) count += 1 print(f"{count} - {email}") mylist = [f for f in glob.glob("*.csv")] for file in mylist: i_file_name = file o_file_name = f"hashed-{file}" print(f"开始处理: {i_file_name}") hash_file(i_file_name, o_file_name)
核心修复说明
- 空文件防护:用
try-except捕获空文件的StopIteration异常,避免直接崩溃。 - 无效行过滤:检查每一行是否为空或列数不足,跳过并打印提示,方便定位问题行。
- 空邮箱处理:对邮箱字符串做去空格处理,遇到空字符串时输出空值保持结构一致。
- 移除冗余操作:删掉手动关闭文件的代码,让
with自动管理文件上下文。 - 增强日志:增加文件处理提示,能快速知道哪个文件出了问题。
内容的提问来源于stack exchange,提问作者user20326377
相关产品推荐
相关产品推荐

