You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何不循环将多列生成的bigram列表批量加入DataFrame

错误原因

你直接将元组组成的bigram_list传给append方法时,pandas会默认将每个二元组的两个元素识别为两列的值,与你目标DataFrame仅1列的结构不匹配,因此抛出错误。且append方法本身性能较差,每次调用都会生成新的DataFrame对象,循环调用会大幅降低运行效率。

解决方法

推荐先将所有列生成的bigram统一收集到Python原生列表中,最后一次性转换为DataFrame,避免多次IO操作,性能远高于逐行/逐列追加,同时兼容性更强:

import pandas as pd

# 初始化空列表存储所有bigram
all_bigrams = []
# 取数据集总列数
col_total = txn_corpus.shape[1]

for col_idx in range(col_total):
    # 取当前列的所有值
    cur_col = txn_corpus.iloc[:, col_idx]
    # 生成当前列的bigram列表
    col_bigram = list(zip(cur_col[:-1], cur_col[1:]))
    # 追加到总列表
    all_bigrams.extend(col_bigram)

# 一次性生成目标DataFrame
txn_corpus_pair = pd.DataFrame({"bigram": all_bigrams})

如果你的pandas版本在0.25及以上,也可以用向量化操作实现无循环的更简洁写法:

txn_corpus_pair = pd.DataFrame({
    "bigram": (txn_corpus.apply(lambda x: list(zip(x[:-1], x[1:])))
               .explode()
               .reset_index(drop=True))
})

内容的提问来源于stack exchange,提问作者raghu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 13:45:03