You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用textnets进行文本网络分析时遇ValueError:不允许负维度

ValueError: negative dimensions are not allowed 问题排查与解决

问题场景

完成文本情感分析后,使用textnets库进行文本网络分析,运行t = tn.Textnet(corpus.tokenized())时触发ValueError: negative dimensions are not allowed错误。

完整代码

import textnets as tn
tn.params["seed"] = 42

df = pd.read_csv(r"C:\Users\User\Downloads\archive\nltk_split.csv")
df.head()

df.drop(['Unnamed: 0.2', 'Unnamed: 0.1', 'Unnamed: 0', 'Unnamed: 15', 'Unnamed: 16', 'id', 'favorite_count', 'created_at', 'retweet_count', 'coordinates', 'score', 'neg', 'neu', 'pos', 'compound','created_at_date'], axis=1, inplace=True)
df.head()

corpus = tn.Corpus.from_df(df, doc_col='text')
t = tn.Textnet(corpus.tokenized())
t.plot(label_nodes=True,
       show_clusters=True)

报错堆栈

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
~\AppData\Local\Temp\ipykernel_19036\2509364554.py in <module>
----> 1 t = tn.Textnet(corpus.tokenized())

~\AppData\Roaming\Python\Python39\site-packages\textnets\network.py in __init__(self, data, min_docs, connected, remove_weak_edges, doc_attrs)
    353             self._matrix = data
    354         elif isinstance(data, (TidyText, pd.DataFrame)):
---> 355             self._matrix = _im_from_tidy_text(data, min_docs)
    356         if remove_weak_edges:
    357             pairs: pd.Series = self._matrix.stack()

~\AppData\Roaming\Python\Python39\site-packages\textnets\network.py in _im_from_tidy_text(tidy_text, min_docs)
    812         .set_index("label")
    813     )
---> 814     im = tt[tt["keep"]].pivot(values="term_weight", columns="term").fillna(0)
    815     return IncidenceMatrix(im)
    816 

E:\anaconda\lib\site-packages\pandas\core\frame.py in pivot(self, index, columns, values)
   7883         from pandas.core.reshape.pivot import pivot
   7884 
-> 7885         return pivot(self, index=index, columns=columns, values=values)
   7886 
   7887     _shared_docs[

E:\anaconda\lib\site-packages\pandas\core\reshape\pivot.py in pivot(data, index, columns, values)
    518         else:
    519             indexed = data._constructor_sliced(data[values]._values, index=multiindex)
---> 520     return indexed.unstack(columns_listlike)
    521 
    522 

E:\anaconda\lib\site-packages\pandas\core\series.py in unstack(self, level, fill_value)
   4155         from pandas.core.reshape.reshape import unstack
   4156 
-> 4157         return unstack(self, level, fill_value)
   4158 
   4159     # ----------------------------------------------------------------------

E:\anaconda\lib\site-packages\pandas\core\reshape\reshape.py in unstack(obj, level, fill_value)
    489         if is_1d_only_ea_dtype(obj.dtype):
    490             return _unstack_extension_series(obj, level, fill_value)
---> 491         unstacker = _Unstacker(
    492             obj.index, level=level, constructor=obj._constructor_expanddim
    493         )

E:\anaconda\lib\site-packages\pandas\core\reshape\reshape.py in __init__(self, index, level, constructor)
    138             )
    139 
---> 140         self._make_selectors()
    141 
    142     @cache_readonly

E:\anaconda\lib\site-packages\pandas\core\reshape\reshape.py in _make_selectors(self)
    186 
    187         selector = self.sorted_labels[-1] + stride * comp_index + self.lift
---> 188         mask = np.zeros(np.prod(self.full_shape), dtype=bool)
    189         mask.put(selector, True)
    190 

ValueError: negative dimensions are not allowed 

错误原因

该错误源于pandas的pivot/unstack操作时生成的数组维度为负数,核心诱因包括:

  • 输入的文本数据存在大量空值或空白内容,分词后无有效术语
  • textnets默认min_docs=2(术语至少出现在2篇文档中),若所有术语仅出现1次,过滤后无有效数据
  • 分词后的tidy文本存在重复的(文档标签, 术语)组合,导致维度计算异常

解决方法

1. 清理文本数据

先检查并移除空值和空白文本:

# 检查text列空值数量
print(df['text'].isnull().sum())
# 移除空值行
df = df.dropna(subset=['text'])
# 移除纯空白文本行
df = df[df['text'].str.strip() != '']

2. 调整min_docs参数

降低术语出现的文档数阈值,比如设为1:

t = tn.Textnet(corpus.tokenized(), min_docs=1)

3. 验证分词结果

确认分词后存在有效术语:

tokenized_corpus = corpus.tokenized()
# 查看分词数据前几行
print(tokenized_corpus.head())
# 检查唯一术语数量
print(f"唯一术语数量: {tokenized_corpus['term'].nunique()}")

4. 处理重复条目

若存在重复的(文档标签, 术语)组合,先去重:

tokenized_corpus = tokenized_corpus.drop_duplicates(subset=['label', 'term'])
t = tn.Textnet(tokenized_corpus)

内容的提问来源于stack exchange,提问作者Karthik Bhandary

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 15:00:12