You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python词关联分析返回空结果问题求助(附代码与数据集)

问题排查与修复方案

核心问题分析

你的代码存在几处语法错误和逻辑缺陷,导致词关联结果为空:

  1. 导入语句语法错误
    原代码里的多模块导入写法不符合Python语法,会直接报错中断执行。
  2. 数据清理行语法错误
    df = df.dropna() df.isna().any() 将两行代码合并成一行,触发语法错误,导致后续流程无法正常运行。
  3. 关键词匹配大小写不统一
    文本已被转成全小写,但用户输入的关键词可能是大写或混合大小写,导致匹配失败。
  4. Bigram筛选逻辑倒置
    先取频率最高的10个Bigram再筛选含关键词的,大概率这10个里没有目标关键词,自然返回空。
  5. 未过滤低频率Bigram
    没有设置最小出现次数,包含大量低频无意义组合,干扰有效结果提取。

修复后的完整代码

import pandas as pd 
import nltk
from nltk.tokenize import word_tokenize
from nltk.collocations import BigramAssocMeasures, BigramCollocationFinder 
from nltk.corpus import stopwords 
import string

# 下载NLTK所需资源(首次运行需要)
nltk.download('punkt')
nltk.download('stopwords')

# 加载CSV数据集
df = pd.read_csv('mtsamples.csv')

# 数据清理:删除空值并验证
df = df.dropna()
print(df.isna().any())  # 验证是否还有空值

# 定义标点移除函数
def remove_punctuation(text):
    translator = str.maketrans('', '', string.punctuation)
    return text.translate(translator)

# 定义停用词移除函数
stop_words = set(stopwords.words('english'))
def remove_stopwords(text):
    tokens = text.split()
    filtered_tokens = [token for token in tokens if token not in stop_words]
    return ' '.join(filtered_tokens)

# 文本预处理流水线
df['transcription'] = df['transcription'].apply(remove_punctuation)
df['transcription'] = df['transcription'].str.lower()
df['transcription'] = df['transcription'].apply(remove_stopwords)

# 获取用户输入并统一为小写
keyword = input('Enter a keyword: ').lower()

# 分词处理
tokenized_text = [word_tokenize(text) for text in df['transcription']]

# 初始化Bigram查找器
finder = BigramCollocationFinder.from_documents(tokenized_text)

# 过滤掉出现次数少于5次的Bigram(可根据数据集大小调整)
finder.apply_freq_filter(5)

# 定义关联度量
measures = BigramAssocMeasures()

# 先筛选包含关键词的Bigram,再按原始频率排序取Top10
matching_bigrams = [bigram for bigram in finder.ngram_fd if keyword in bigram]
# 按频率排序
matching_bigrams_sorted = sorted(matching_bigrams, key=lambda x: finder.ngram_fd[x], reverse=True)[:10]

# 输出结果
print('Frequent bigrams:')
if matching_bigrams_sorted:
    for bigram in matching_bigrams_sorted:
        print(' '.join(bigram))
else:
    print(f"No bigrams found containing keyword '{keyword}'")

关键修复点说明

  • 修正导入语法:将合并的导入语句拆分为独立行,确保模块正常导入。
  • 修复数据清理代码:拆分合并的两行代码,保证空值删除操作正确执行。
  • 统一关键词大小写:将用户输入的关键词转为小写,和预处理后的文本保持一致,避免匹配失败。
  • 调整筛选逻辑:先从所有Bigram中筛选含关键词的,再按频率排序取Top10,确保不会错过包含目标词的组合。
  • 添加频率过滤:用apply_freq_filter(5)过滤掉出现次数少于5次的Bigram,减少噪音,聚焦有效关联。

内容的提问来源于stack exchange,提问作者PeatyBoWeaty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 12:47:25