You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python遍历含文本的行生成Bigrams并关联原DataFrame表格

解决方法

你的函数问题分析

  • 传入函数的是字符串而非列表,直接用len(input_list)会计算字符长度,不是单词数量,必须先把文本分割成单词列表。
  • 循环内append(input_list[1:])逻辑错误,这是截取子列表,不是生成相邻单词的二元组。
  • return语句写在循环内部,第一次循环就直接返回,且当输入文本只有单个单词时,函数无返回值,默认返回None,这就是你得到全None的原因。

正确实现方式

方法1:修正自定义函数

def find_bigrams(text):
    # 去除首尾空格并分割成单词列表
    words = text.strip().split()
    bigram_list = []
    # 遍历生成相邻单词的二元组
    for i in range(len(words) - 1):
        bigram_list.append((words[i], words[i+1]))
    return bigram_list

应用到DataFrame:

df['Content'] = df['Content'].apply(find_bigrams)

方法2:用列表推导式简化

df['Content'] = df['Content'].apply(
    lambda x: [(x.split()[i], x.split()[i+1]) for i in range(len(x.split())-1)] if len(x.split()) >=2 else []
)

执行后,Content列会生成你期望的二元组列表,同时保留原表格中Company、Code等列的关联关系。

内容的提问来源于stack exchange,提问作者Dhanya_mj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 11:40:14