Unicode 16.0与15.1字符生成及新增字符提取问题求助
Unicode 16.0与15.1字符生成及新增字符提取问题求助
我正在尝试实现以下需求:
- 生成一个包含所有Unicode 15.1字符的文件
- 生成一个包含所有Unicode 16.0字符的文件
- 生成第三个文件,专门展示Unicode 16.0相比15.1新增的字符
我写了一段代码尝试实现,但结果并不符合预期。主要有两个疑问:
- 可能存在一些在Unicode 15.1中不可打印,但在Unicode 16.0中可打印的新表情或其他字符,我的代码没有考虑到这种情况
- 我不确定自己的字符生成逻辑是否正确
麻烦各位帮忙看看我的源代码,指出问题所在,谢谢!
import os file_15_1 = "unicode_15_1.txt" file_16_0 = "unicode_16_0.txt" file_new_in_16_0 = "new_in_16_0.txt" unicode_15_1_end = 149813 unicode_16_0_end = 154998 def is_visible(char): return char.isprintable() and not char.isspace() and char != "" def generate_unicode_file(start, end, filename): with open(filename, "w", encoding="utf-8") as f: for codepoint in range(start, end + 1): try: f.write(chr(codepoint) + "\n") except ValueError: continue generate_unicode_file(0, unicode_15_1_end, file_15_1) generate_unicode_file(0, unicode_16_0_end, file_16_0) def find_new_characters(file1, file2, output_file): with open(file1, "r", encoding="utf-8") as f1, open(file2, "r", encoding="utf-8") as f2: chars_15_1 = set(f1.read().splitlines()) chars_16_0 = set(f2.read().splitlines()) new_in_16_0 = chars_16_0 - chars_15_1 # No with open(output_file, "w", encoding="utf-8") as f_out: for char in sorted(new_in_16_0): if is_visible(char): f_out.write(char + "\n") find_new_characters(file_15_1, file_16_0, file_new_in_16_0) print(f"- {file_15_1}") print(f"- {file_16_0}") print(f"- {file_new_in_16_0}")
备注:内容来源于stack exchange,提问作者Silme94
相关产品推荐
相关产品推荐

