You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将TensorFlow张量转为NumPy数组并将my_li列表转换为该格式

问题描述

处理数据时,df['name']里全是字符串,我跑了这段代码:

a = list(df['name'])

text_vectorization = TextVectorization(output_mode = "int",
                                       max_tokens=100,
                                       output_length=10)


my_li = []
for i in a:
  my_li.append(list(text_vectorization(i)))

my_li[0]

得到的my_li[0]是一堆tf.Tensor:

[<tf.Tensor: shape=(), dtype=int64, numpy=22>, <tf.Tensor: shape=(), dtype=int64, numpy=79>, <tf.Tensor: shape=(), dtype=int64, numpy=13>, <tf.Tensor: shape=(), dtype=int64, numpy=3>, <tf.Tensor: shape=(), dtype=int64, numpy=8>, <tf.Tensor: shape=(), dtype=int64, numpy=4>, <tf.Tensor: shape=(), dtype=int64, numpy=5>, <tf.Tensor: shape=(), dtype=int64, numpy=63>, <tf.Tensor: shape=(), dtype=int64, numpy=34>, <tf.Tensor: shape=(), dtype=int64, numpy=2>, <tf.Tensor: shape=(), dtype=int64, numpy=10>, <tf.Tensor: shape=(), dtype=int64, numpy=25>, <tf.Tensor: shape=(), dtype=int64, numpy=74>, <tf.Tensor: shape=(), dtype=int64, numpy=53>, <tf.Tensor: shape=(), dtype=int64, numpy=85>, <tf.Tensor: shape=(), dtype=int64, numpy=7>, <tf.Tensor: shape=(), dtype=int64, numpy=27>, <tf.Tensor: shape=(), dtype=int64, numpy=54>, <tf.Tensor: shape=(), dtype=int64, numpy=125>, <tf.Tensor: shape=(), dtype=int64, numpy=76>]

现在要把这些Tensor转成NumPy值,再把整个my_li变成NumPy数组,用来喂模型训练,该咋弄?

解决方案

方法1:生成列表时直接转成numpy值

要是还没生成my_li,直接在循环里就处理好,省得后续二次操作:

a = list(df['name'])

text_vectorization = TextVectorization(output_mode = "int",
                                       max_tokens=100,
                                       output_length=10)

my_li = []
for i in a:
    # 先把Tensor转成numpy数组,再转成列表存起来
    vec = text_vectorization(i).numpy().tolist()
    my_li.append(vec)

# 最后转成NumPy数组
my_np_array = np.array(my_li)

方法2:给已有的my_li批量转值

要是已经有了那个全是Tensor的my_li,用嵌套列表推导式提取每个Tensor的numpy值,再转成数组:

import numpy as np

# 遍历每个子列表,把每个Tensor转成numpy值
processed_li = [[tensor.numpy() for tensor in sublist] for sublist in my_li]
# 转成最终的NumPy数组
my_np_array = np.array(processed_li)

方法3:直接批量处理(最推荐)

其实TextVectorization支持直接喂整个字符串列表,不用循环一个个处理,效率高多了:

import numpy as np

# 直接批量处理所有字符串
vectorized_tensor = text_vectorization(a)
# 转成NumPy数组
my_np_array = vectorized_tensor.numpy()

处理完得到的my_np_array是(样本数量, output_length)形状的数组,直接就能给模型训练用。

内容的提问来源于stack exchange,提问作者gfd_tt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 21:05:16