如何用Python移除文本文件中的^M(Control M)字符并保持行连续
移除文件中Control M字符并保持行结构的解决方案
问题描述
我的原始文件内容如下:
This line has control character ^M this is bad I will try it
我想用Python移除文件中的Control M字符(即\r),生成的目标文件内容应该是:
This line has control character this is bad I will try it
我尝试了两种替换方式:
line.replace("\r", "")line.replace("\r\n", "")
使用的代码片段如下:
with open(file_path, "r") as input_file: lines = input_file.readlines() new_lines = [] for line in lines: new_line = line.replace("\r", "") new_lines.append(new_line) new_file_name = "replace_control_char.dat" new_file_path = os.path.join(here, data_dir, new_file_name) with open(new_file_path, "w") as output_file: for line in new_lines: output_file.write(line)
但生成的新文件结果却变成了这样:
This line has control character this is bad I will try it
原本同一行的内容被拆成了两行,我需要移除Control M后仍保持内容在同一行,求解决办法。
原因分析
问题出在readlines()的默认处理逻辑上——Python用文本模式读文件时,会把\r、\n、\r\n都当成换行符,所以原本带\r的单行内容被拆成了两行读进来。之后替换掉\r,但已经拆好的两行没法自动合并,就出现了现在的问题。
解决方法
方法一:二进制模式读写
直接用二进制模式读取整个文件,替换掉\r对应的字节后再写入,完全保留原始行结构:
import os with open(file_path, "rb") as input_file: content = input_file.read() # 替换\r字节为空字节 new_content = content.replace(b"\r", b"") new_file_name = "replace_control_char.dat" new_file_path = os.path.join(here, data_dir, new_file_name) with open(new_file_path, "wb") as output_file: output_file.write(new_content)
方法二:指定newline=''读取
在文本模式下打开文件时加上newline=''参数,Python就不会自动转换行结束符,能完整读取原始行内容,再替换掉行内的\r即可:
import os with open(file_path, "r", newline='') as input_file: lines = input_file.readlines() new_lines = [] for line in lines: # 替换行内的\r字符 new_line = line.replace("\r", "") new_lines.append(new_line) new_file_name = "replace_control_char.dat" new_file_path = os.path.join(here, data_dir, new_file_name) with open(new_file_path, "w") as output_file: for line in new_lines: output_file.write(line)
这两种方法都能确保移除Control M字符后,原本的单行内容不会被拆分。
内容的提问来源于stack exchange,提问作者Arthur
相关产品推荐
相关产品推荐

