Python逐行读取TXT文件时部分行内容反向的技术咨询
TXT文件读取部分行内容反向的原因与解决方法
问题描述
尝试用以下Python代码读取路径C:/Users/Lenovo/Documents/LinuxLab/file.txt的书籍元数据:
f = open("C:/Users/Lenovo/Documents/LinuxLab/file.txt",mode = 'r',encoding='utf-8') for line in f: print(line)
待读取的TXT文件内容:
Title : Linux System Programming Talking Directly to the Kernel and C Library Publisher : O'Reilly Media Edition : 2 Year : 2013 Month : 1 Language : English Paperback : 456 pages ISBN-10 : 1449339530 ISBN-13 : 978-1449339531
执行后部分行内容反向,输出如下:
Title : Linux System Programming Talking Directly to the Kernel and C Library
O'Reilly Media : Publisher
Edition : 2
Year : 2013
Month : 1
English : Language
456 pages : Paperback
1449339530 : ISBN-10
978-1449339531: ISBN-13
原因
问题出在TXT文件中隐藏的双向文本控制字符(具体是U+200F从右到左标记,即你看到的)。这类字符会强制改变文本的显示方向,当它们出现在行内时,会让后续内容从右到左渲染,导致键值对顺序被反转。
解决方法
读取文件时过滤掉所有双向控制字符即可,以下是两种可行方案:
方案1:用unicodedata过滤所有控制类字符
import unicodedata def clean_control_chars(text): # 移除所有Unicode控制类字符 return ''.join(c for c in text if unicodedata.category(c)[0] != 'C') with open("C:/Users/Lenovo/Documents/LinuxLab/file.txt", mode='r', encoding='utf-8') as f: for line in f: cleaned_line = clean_control_chars(line).strip() print(cleaned_line)
方案2:直接替换特定双向控制符
如果确认只有U+200F(从右到左标记)和U+200E(从左到右标记)导致问题,可直接替换:
with open("C:/Users/Lenovo/Documents/LinuxLab/file.txt", mode='r', encoding='utf-8') as f: for line in f: cleaned_line = line.replace('\u200F', '').replace('\u200E', '').strip() print(cleaned_line)
两种方法都能清理掉导致方向反转的隐藏字符,让输出恢复正常的键值顺序。
内容的提问来源于stack exchange,提问作者Tareq Ewaida
相关产品推荐
相关产品推荐

