Python csv模块写入CSV列数据错位问题求助(禁用Pandas)
Python写入CSV时B列数据偏移至A列的解决方法
问题背景
- 用Python将含英文逗号、数字、阿拉伯文本的UTF-8数据写入CSV文件时,部分B列(
Content列)数据被错误写入A列(File Name列)。 - 记事本中查看数据已被双引号包裹,显示正常;但MS Office和LibreOffice预览时列偏移异常,LibreOffice打开文件后显示正常。
- 无法使用Pandas,需保持文件打开状态分批写入数据。
- 添加特定数据后,文件被识别为「Unicode text, UTF-8 text, with CRLF, LF line terminators」而非「CSV text」,当前使用第一个代码片段。
已尝试的代码
# 代码片段1 with open(df_path, "w", newline="", encoding="utf-8") as csv_file: writer = csv.DictWriter(csv_file, fieldnames=["File Name", "Content"], quoting=csv.QUOTE_ALL) writer.writeheader() writer.writerow({"File Name": file, "Content": txt})
# 代码片段2 with open(df_path, "w", newline="", encoding="utf-8") as csv_file: writer = csv.writer(csv_file) writer.writerow(["File Name", "Content"]) writer.writerow([file, '"' + txt + '"'])
# 代码片段3 with open(df_path, "w", newline="", encoding="utf-8") as csv_file: writer = csv.DictWriter(csv_file, fieldnames=["File Name", "Content"]) writer.writeheader() writer.writerow({"File Name": file, "Content": txt})
# 代码片段4 with open(df_path, "w", newline="", encoding="utf-8") as csv_file: writer = csv.DictWriter(csv_file, fieldnames=["File Name", "Content"], quoting=csv.QUOTE_ALL) writer.writeheader() writer.writerow({"File Name": file, "Content": txt})
# 代码片段5 with open(df_path, "w", newline="", encoding="utf-8") as csv_file: writer = csv.DictWriter(csv_file, fieldnames=["File Name", "Content"], delimiter=",") writer.writeheader() writer.writerow({"File Name": file, "Content": txt})
解决方案
1. 清理文本中的换行/回车符
CSV解析器会将内容中的\n、\r误判为行分隔符,导致列偏移。写入前先清理这些字符:
# 替换换行和回车为空格,或其他不影响的分隔符 cleaned_txt = txt.replace('\n', ' ').replace('\r', '')
2. 使用UTF-8 BOM编码并配置转义符
MS Office对无BOM的UTF-8 CSV兼容性差,改用utf-8-sig编码;同时设置escapechar确保双引号被正确转义,避免解析混乱:
import csv with open(df_path, "w", newline="\r\n", encoding="utf-8-sig") as csv_file: # QUOTE_ALL确保所有字段被引号包裹,escapechar处理字段内的引号 writer = csv.DictWriter( csv_file, fieldnames=["File Name", "Content"], quoting=csv.QUOTE_ALL, escapechar='\\' ) writer.writeheader() writer.writerow({"File Name": file, "Content": cleaned_txt})
3. 统一行终止符
文件同时存在CRLF和LF会导致解析错误,强制使用Windows标准的CRLF行终止符:
在打开文件时指定newline="\r\n",如上面代码所示。
4. 保持文件打开状态的分批写入
若需要持续写入数据,不要反复打开/关闭文件,保持writer对象复用:
import csv # 初始化打开文件,保持写入状态 csv_file = open(df_path, "w", newline="\r\n", encoding="utf-8-sig") writer = csv.DictWriter( csv_file, fieldnames=["File Name", "Content"], quoting=csv.QUOTE_ALL, escapechar='\\' ) writer.writeheader() # 模拟分批写入数据,替换成你的实际数据迭代逻辑 for file, txt in [("file1.txt", "内容1,含逗号"), ("file2.txt", "阿拉伯文本: مرحبا")]: cleaned_txt = txt.replace('\n', ' ').replace('\r', '') writer.writerow({"File Name": file, "Content": cleaned_txt}) # 所有数据写入完成后再关闭文件 csv_file.close()
5. 验证CSV格式正确性
写入后可以用Python的csv.reader验证列数是否正常:
with open(df_path, "r", encoding="utf-8-sig") as f: reader = csv.reader(f) for idx, row in enumerate(reader): print(f"第{idx+1}行列数: {len(row)}") # 正常情况下每行应该是2列
内容的提问来源于stack exchange,提问作者Esraa Abdelmaksoud
相关产品推荐
相关产品推荐

