You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取Unicode 13.0.0全部143859个字符列表并在Python中打印

解决方案

Unicode 13.0.0的全量收录字符无需额外获取外部列表,直接通过Python筛选有效码点即可得到,官方统计的143859个字符指已分配、有正式名称的字符,不含私有使用区、未分配码点和无名称控制字符。

方案1:使用Python标准库unicodedata

要求当前Python内置的Unicode版本与13.0.0匹配,先执行以下代码校验版本:

import unicodedata
print(unicodedata.unidata_version)

若输出为13.0.0,可直接使用以下代码获取全量字符:

import unicodedata

assert unicodedata.unidata_version == "13.0.0", "Unicode版本不匹配,请使用方案2"

unicode_13_chars = []
# 遍历所有Unicode有效码点范围 0x0000 ~ 0x10FFFF
for codepoint in range(0x110000):
    try:
        char = chr(codepoint)
        # 能获取到正式名称的即为官方收录的统计字符
        char_name = unicodedata.name(char)
        unicode_13_chars.append((codepoint, char, char_name))
    except ValueError:
        # 过滤无名称的控制字符、未分配码点、私有区字符
        continue

# 验证数量
print(f"共获取Unicode 13.0.0收录字符:{len(unicode_13_chars)}个")

# 打印所有字符示例(注释解除即可运行,注意打印大量生僻字符可能导致终端卡顿)
# for cp, char, name in unicode_13_chars:
#     print(f"U+{cp:04X} | {char} | {name}")

# 打印前20个字符示例
print("前20个字符预览:")
for cp, char, name in unicode_13_chars[:20]:
    print(f"U+{cp:04X} | {char} | {name}")

方案2:使用第三方库固定Unicode版本

如果内置unicodedata版本不匹配,可使用指定版本的unicodedata2库获取对应版本字符:

  1. 安装对应版本库:
    pip install unicodedata2==13.0.0.post2
  2. 执行以下代码获取字符:
import unicodedata2

assert unicodedata2.unidata_version == "13.0.0"

unicode_13_chars = []
for codepoint in range(0x110000):
    try:
        char = chr(codepoint)
        char_name = unicodedata2.name(char)
        unicode_13_chars.append((codepoint, char, char_name))
    except ValueError:
        continue

print(f"共获取Unicode 13.0.0收录字符:{len(unicode_13_chars)}个")

注意事项

  • 若需要保留控制字符、私有使用区字符,删除try-except过滤逻辑直接收集chr(codepoint)即可,最终收集到的字符数会远大于143859
  • 终端打印生僻字符、特殊符号、emoji时如果出现乱码,属于终端字体或渲染能力不足,不影响字符本身获取的正确性
  • 打印全量字符前建议先存储到文件,避免大量字符输出导致终端卡顿崩溃

内容的提问来源于stack exchange,提问作者nimish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 19:15:01