如何将Python 2输出的Unicode编码波斯文转换为实际字符?
解决Python2生成的Unicode转义字符串转波斯文实际字符的问题
我刚好处理过类似的场景,给你几个靠谱的解决办法,帮你把那些u'\u0641\u0632\u0646\u062f\u0627\u0646'形式的转义字符串转换成实际的波斯文字符:
方法1:批量转换整个文件(推荐)
写个简单的Python脚本,一次性把整个文件的转义内容转换成实际字符。这个方法用ast.literal_eval来安全解析字符串,比直接用eval更稳妥,不会执行恶意代码:
import ast def convert_vocab_unicode(input_path, output_path): # 读取原文件内容 with open(input_path, 'r', encoding='utf-8') as infile: raw_content = infile.read() # 先去掉Python2风格的u前缀,再用ast解析转义字符 processed_content = raw_content.replace("u'", "'") converted_content = ast.literal_eval(f'"{processed_content}"') # 写入转换后的文件 with open(output_path, 'w', encoding='utf-8') as outfile: outfile.write(converted_content) # 替换成你的文件路径 convert_vocab_unicode('original_vocab.txt', 'converted_vocab.txt')
方法2:读取时动态转换
如果你的后续流程是用Python读取词汇文件并实时比对,可以在读取每行的时候直接转换,不用提前修改文件:
import ast with open('original_vocab.txt', 'r', encoding='utf-8') as vocab_file: for line in vocab_file: # 清理每行内容,去掉u前缀并解析转义 cleaned_line = line.strip().replace("u'", "'") actual_persian_text = ast.literal_eval(f'"{cleaned_line}"') # 这里就可以用actual_persian_text和你的目标词汇做比对了 print(actual_persian_text) # 输出实际的波斯文字符
方法3:命令行快速处理(适合小文件)
如果文件不大,不想写脚本,直接用Python的命令行参数快速转换:
python -c "import ast; content = open('original_vocab.txt').read().replace('u\\'', '\''); print(ast.literal_eval(f'\"{content}\"'))" > converted_vocab.txt
特殊情况处理:混合内容的文件
如果你的词汇文件里不只有u'...'格式的字符串,还有其他内容,可以用正则匹配精准替换每个转义字符串:
import re import ast def convert_target_matches(input_path, output_path): with open(input_path, 'r', encoding='utf-8') as infile: content = infile.read() # 定义匹配u'xxx'格式的正则 unicode_pattern = r"u'([^']*)'" # 替换每个匹配项 def replace_unicode(match): # 解析捕获到的转义字符串 return ast.literal_eval(f'"{match.group(1)}"') converted_content = re.sub(unicode_pattern, replace_unicode, content) with open(output_path, 'w', encoding='utf-8') as outfile: outfile.write(converted_content)
这些方法都能把Unicode转义序列转换成实际的波斯文字符,你可以根据自己的场景选最合适的。
内容的提问来源于stack exchange,提问作者Gmosy Gnaq
相关产品推荐
相关产品推荐

