如何在TensorFlow中不使用py_func实现Python的ord函数功能
解决字符串逐字符转ASCII十进制数值(不使用py_func)
我懂你碰到的痛点——要在需要保存/恢复模型的场景下,把字符串占位符里的每个字符转成对应的ASCII十进制值,但是py_func会破坏模型的可序列化性,确实得找框架原生的方案。如果是用TensorFlow的话,完全可以用内置的字符串操作来实现,不需要依赖Python自定义函数。
核心思路
用TensorFlow的原生字符串处理API,分两步完成:
- 把输入字符串拆分成单个字符的张量
- 将每个字符解码为对应的ASCII码点(也就是十进制数值)
TensorFlow 2.x 实现示例
import tensorflow as tf # 定义一个可序列化的函数,用于字符串转ASCII @tf.function def string_to_ascii_tensor(input_str): # 1. 将每个输入字符串拆分为单个字符的张量 split_characters = tf.strings.unicode_split(input_str, input_encoding='ASCII') # 2. 解码字符得到ASCII十进制数值 ascii_values = tf.strings.unicode_decode(split_characters, input_encoding='ASCII') return ascii_values # 测试使用 if __name__ == "__main__": test_input = tf.constant(["abc123", "XYZ!@#"]) result = string_to_ascii_tensor(test_input) print(result.numpy()) # 输出:[[ 97 98 99 49 50 51] # [120 121 122 33 64 35]]
TensorFlow 1.x 兼容版本
如果还在使用TF1.x的占位符模式,代码可以这样写:
import tensorflow as tf # 定义字符串占位符 str_placeholder = tf.compat.v1.placeholder(tf.string, shape=[None]) # 拆分字符并转ASCII split_chars = tf.strings.unicode_split(str_placeholder, input_encoding='ASCII') ascii_values = tf.strings.unicode_decode(split_chars, input_encoding='ASCII') # 测试会话 with tf.compat.v1.Session() as sess: test_strings = ["hello", "world"] output = sess.run(ascii_values, feed_dict={str_placeholder: test_strings}) print(output) # 输出:[[104 101 108 108 111] # [119 111 114 108 100]]
额外注意事项
- 如果输入字符串可能包含非ASCII字符,可以提前过滤,避免解码报错:
# 过滤掉所有非ASCII字符 filtered_str = tf.strings.regex_replace(input_str, r'[^\x00-\x7F]', '') - 这些操作都是TensorFlow的原生Op,完全支持模型的保存和恢复,不会像
py_func那样引入不可序列化的逻辑。
内容的提问来源于stack exchange,提问作者Nicola Donadoni
相关产品推荐
相关产品推荐

