You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Unicode 16.0与15.1字符生成及新增字符提取问题求助

Unicode 16.0与15.1字符生成及新增字符提取问题求助

我正在尝试实现以下需求:

  • 生成一个包含所有Unicode 15.1字符的文件
  • 生成一个包含所有Unicode 16.0字符的文件
  • 生成第三个文件,专门展示Unicode 16.0相比15.1新增的字符

我写了一段代码尝试实现,但结果并不符合预期。主要有两个疑问:

  1. 可能存在一些在Unicode 15.1中不可打印,但在Unicode 16.0中可打印的新表情或其他字符,我的代码没有考虑到这种情况
  2. 我不确定自己的字符生成逻辑是否正确

麻烦各位帮忙看看我的源代码,指出问题所在,谢谢!

import os

file_15_1 = "unicode_15_1.txt"
file_16_0 = "unicode_16_0.txt"
file_new_in_16_0 = "new_in_16_0.txt"

unicode_15_1_end = 149813
unicode_16_0_end = 154998 

def is_visible(char):
    return char.isprintable() and not char.isspace() and char != ""

def generate_unicode_file(start, end, filename):
    with open(filename, "w", encoding="utf-8") as f:
        for codepoint in range(start, end + 1):
            try:
                f.write(chr(codepoint) + "\n")
            except ValueError:
                continue

generate_unicode_file(0, unicode_15_1_end, file_15_1)
generate_unicode_file(0, unicode_16_0_end, file_16_0)

def find_new_characters(file1, file2, output_file):
    with open(file1, "r", encoding="utf-8") as f1, open(file2, "r", encoding="utf-8") as f2:
        chars_15_1 = set(f1.read().splitlines())  
        chars_16_0 = set(f2.read().splitlines())  
        new_in_16_0 = chars_16_0 - chars_15_1 # No

    with open(output_file, "w", encoding="utf-8") as f_out:
        for char in sorted(new_in_16_0):
            if is_visible(char): 
                f_out.write(char + "\n")


find_new_characters(file_15_1, file_16_0, file_new_in_16_0)

print(f"- {file_15_1}")
print(f"- {file_16_0}")
print(f"- {file_new_in_16_0}")

备注:内容来源于stack exchange,提问作者Silme94

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 14:42:58