You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何通过f-string或其他方法将<U+XXXX>编码转为实际Unicode字符

问题原因

Python的\u转义序列是语法层面的硬规则,只能在代码编写阶段直接跟4位十六进制字符,不能通过字符串拼接、f-string动态注入内容。你写的f'\u{a[3:7]}'会在Python解析字符串阶段就检测到\u后面没有跟合法的4位十六进制数,直接抛出语法错误,等不到f-string的变量替换逻辑执行。

解决方案

用Python内置的chr()函数实现需求:先把提取到的十六进制编码转成十进制整数,再通过chr()生成对应的Unicode字符即可。

优化后的代码写法

写法1:保留原有遍历逻辑

import re

def replace_unicode(s):
    # 正则直接用分组提取4位编码,无需后续切片处理
    uni = re.findall(r'<U\+(\w{4})>', s)
    for code in uni:
        s = s.replace(f'<U+{code}>', chr(int(code, 16)))
    return s

写法2:用re.sub回调实现一步替换(更简洁)

不需要先找所有匹配再遍历替换,直接通过re.sub的回调函数处理每一个匹配项:

import re

def replace_unicode(s):
    return re.sub(r'<U\+(\w{4})>', lambda match: chr(int(match.group(1), 16)), s)
测试验证

用示例字符串测试:

a = "testing test<U+00FA>ing <U+00F3>"
print(replace_unicode(a))
# 输出:testing testúing ó

内容的提问来源于stack exchange,提问作者Kemian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 12:27:05