You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在TensorFlow中不使用py_func实现Python的ord函数功能

解决字符串逐字符转ASCII十进制数值(不使用py_func)

我懂你碰到的痛点——要在需要保存/恢复模型的场景下,把字符串占位符里的每个字符转成对应的ASCII十进制值,但是py_func会破坏模型的可序列化性,确实得找框架原生的方案。如果是用TensorFlow的话,完全可以用内置的字符串操作来实现,不需要依赖Python自定义函数。

核心思路

用TensorFlow的原生字符串处理API,分两步完成:

  • 把输入字符串拆分成单个字符的张量
  • 将每个字符解码为对应的ASCII码点(也就是十进制数值)

TensorFlow 2.x 实现示例

import tensorflow as tf

# 定义一个可序列化的函数,用于字符串转ASCII
@tf.function
def string_to_ascii_tensor(input_str):
    # 1. 将每个输入字符串拆分为单个字符的张量
    split_characters = tf.strings.unicode_split(input_str, input_encoding='ASCII')
    # 2. 解码字符得到ASCII十进制数值
    ascii_values = tf.strings.unicode_decode(split_characters, input_encoding='ASCII')
    return ascii_values

# 测试使用
if __name__ == "__main__":
    test_input = tf.constant(["abc123", "XYZ!@#"])
    result = string_to_ascii_tensor(test_input)
    print(result.numpy())
    # 输出:[[ 97  98  99  49  50  51]
    #        [120 121 122  33  64  35]]

TensorFlow 1.x 兼容版本

如果还在使用TF1.x的占位符模式,代码可以这样写:

import tensorflow as tf

# 定义字符串占位符
str_placeholder = tf.compat.v1.placeholder(tf.string, shape=[None])

# 拆分字符并转ASCII
split_chars = tf.strings.unicode_split(str_placeholder, input_encoding='ASCII')
ascii_values = tf.strings.unicode_decode(split_chars, input_encoding='ASCII')

# 测试会话
with tf.compat.v1.Session() as sess:
    test_strings = ["hello", "world"]
    output = sess.run(ascii_values, feed_dict={str_placeholder: test_strings})
    print(output)
    # 输出:[[104 101 108 108 111]
    #        [119 111 114 108 100]]

额外注意事项

  • 如果输入字符串可能包含非ASCII字符,可以提前过滤,避免解码报错:
    # 过滤掉所有非ASCII字符
    filtered_str = tf.strings.regex_replace(input_str, r'[^\x00-\x7F]', '')
    
  • 这些操作都是TensorFlow的原生Op,完全支持模型的保存和恢复,不会像py_func那样引入不可序列化的逻辑。

内容的提问来源于stack exchange,提问作者Nicola Donadoni

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:03:11