Python中len()获取文本长度返回2n+2值的问题求助
解决Python读取文本文件时len()返回2n+2的问题
跟着《笨办法学Python》练习时,用
len()获取文本文件字符数总是返回2n+2(n为实际字符数),确认文件无多余空行。示例文件内容:
The poem dedicated to Puxijn The Chonk one读取后输出的字符显示为:
ÿþT h e p o e m d e d i c a t e d t o P u x i j n T h e C h o n k o n e怀疑是编码问题,使用最新版Python,求解释原因和解决办法。
测试代码
from sys import argv script, from_file, to_file = argv # 尝试简化命令,结果同样返回2n+2 trial = open(from_file) trial_data = trial.read() print(len(trial_data)) trial.close() # 定义变量后的实际代码 in_file = open(from_file).read() input(f"Transfering {len(in_file)} characters from {from_file} to {to_file}, hit RETURN to continue, CRTL-C to abort.") #'in_data = in_file.read() out_file = open(to_file, 'w').write(in_file)
原因分析
这确实是编码问题,你的文件是用**UTF-16 LE(小端)**编码保存的:
- 开头的
ÿþ是UTF-16 LE的BOM(字节顺序标记),占2个字节,对应你看到的2n+2里的+2; - UTF-16编码中每个ASCII字符都会被存储为2个字节(比如字母
T会变成两个字节),所以实际字符数n对应的字节数就是2n,加上BOM的2字节,总长度就是2n+2,和你看到的结果完全匹配。
解决办法
读取文件时指定正确的编码即可,把open(from_file)改成open(from_file, encoding='utf-16'),Python会自动处理BOM并正确解析字符:
修改后的核心代码片段:
# 简化测试部分 trial = open(from_file, encoding='utf-16') trial_data = trial.read() print(len(trial_data)) # 现在会输出实际字符数 trial.close() # 实际传输部分 in_file = open(from_file, encoding='utf-16').read()
如果需要将文件转存为更通用的UTF-8编码,写入时也可以指定编码:
out_file = open(to_file, 'w', encoding='utf-8').write(in_file)
内容的提问来源于stack exchange,提问作者Mister Mace
相关产品推荐
相关产品推荐

