You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

向形状为(500,151296)的NumPy二维数组添加元素报错求助

问题描述

我有一个形状为(500, 151296)的float32类型NumPy数组,格式如下:

array([[-0.18510018,  0.13180602,  0.32903048, ...,  0.39744213,
        -0.01461623,  0.06420607],
       [-0.14988784,  0.12030973,  0.34801325, ...,  0.36962894,
         0.04133283,  0.04434045],
       [-0.3080041 ,  0.18728344,  0.36068922, ...,  0.09335024,
        -0.11459247,  0.10187756],
       ...,
       [-0.17399777, -0.02492459, -0.07236133, ...,  0.08901921,
        -0.17250113,  0.22222663],
       [-0.17399777, -0.02492459, -0.07236133, ...,  0.08901921,
        -0.17250113,  0.22222663],
       [-0.17399777, -0.02492459, -0.07236133, ...,  0.08901921,
        -0.17250113,  0.22222663]], dtype=float32)

其中单一行array[0]的格式为:

array([-0.18510018,  0.13180602,  0.32903048, ...,  0.39744213,
       -0.01461623,  0.06420607], dtype=float32)

另外我还有一个长度为500的停用词列表:

stopwords = ['no', 'not', 'in' .........]

我想给这个NumPy数组的每一行末尾添加对应的停用词,用了以下代码:

for i in range(len(stopwords)):
  array = np.append(array[i], str(stopwords[i]))

但运行后触发了如下错误:

IndexError                                Traceback (most recent call last)
<ipython-input-45-361e2cf6519b> in <module>
      1 for i in range(len(stopwords)):
----> 2   array = np.append(array[i], str(stopwords[i]))

IndexError: index 2 is out of bounds for axis 0 with size 2

期望输出的array[0]格式为:

array([-0.18510018,  0.13180602,  0.32903048, ...,  0.39744213,
       -0.01461623,  0.06420607, 'no'], dtype=float32)

请问我哪里出错了?


问题分析与解决

错误原因

你的代码存在两个核心问题:

  1. 数组结构被错误覆盖:第一次循环时,np.append(array[i], str(stopwords[i]))会把原数组的第i行(一维数组)和字符串拼接,返回一个新的一维数组并赋值给array,此时原二维数组被替换成了一维数组。第二次循环时,array[i]访问的是这个一维数组的第i个元素(单个float值),再append字符串后得到长度为2的一维数组,第三次循环i=2时,这个数组长度仅为2,访问array[2]自然触发索引越界。
  2. 数据类型冲突:原数组是float32类型,NumPy无法在同一基础类型数组中同时存放数值和字符串,最终数组会自动转为object类型,你期望的输出格式实际无法保持float32 dtype。

正确实现方式

方法1:批量拼接(高效推荐)

把停用词转为二维数组后,用np.column_stack批量拼接:

import numpy as np

# 将停用词转为形状为(500,1)的object类型二维数组
stopwords_arr = np.array(stopwords, dtype=object).reshape(-1, 1)
# 拼接原数组和停用词数组
new_array = np.column_stack((your_original_array, stopwords_arr))

方法2:循环收集结果

如果偏好循环实现,不要覆盖原数组,而是收集处理后的行再转为数组:

new_rows = []
for row, word in zip(your_original_array, stopwords):
    # 将行转为object类型后添加字符串
    new_row = np.append(row.astype(object), word)
    new_rows.append(new_row)
new_array = np.array(new_rows)

说明

  • 由于要同时存储数值和字符串,最终数组的dtype必然是object,无法保持原有的float32类型。
  • 批量操作的效率远高于循环,数组规模较大时优先选择np.column_stack这类方法。

内容的提问来源于stack exchange,提问作者merkle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 23:20:36