You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python词频统计结果失真,求问题排查与优化方案

问题

作为NLP初学者,我有一份超20000行的产品评论数据表(示例如下),已编写Python代码实现评论的词频及短语频次统计并导出CSV。为避免运行卡顿设置了top_n=100,但统计结果严重失真:比如原数据中‘customer service’提及次数远多于统计结果,‘customer’计数仅为9次。需要排查问题原因并提供代码优化方案,同时要关联Product_ID与record_id适配工作工具。

产品评论表示例

Record_IDProduct_IDReview Comment
123489847457I love this product it was shipped fast and is comfortable

当前代码实现

import pandas as pd
from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS
import string
from collections import Counter
from nltk.util import ngrams
import nltk
nltk.download('punkt')

df = pd.read_excel('productsvydata.xlsx')

def preprocess_text(text):
    translator = str.maketrans('', '', string.punctuation)
    text = text.lower() 
    text = text.translate(translator)
    return text

word_counts = {}
phrase_counts = {}

unique_product_ids = df["Product_ID"].unique()

# Set the number of top words and phrases you want to keep
top_n = 100

for selected_product_id in unique_product_ids:
    selected_comments_df = df[df["Product_ID"] == selected_product_id]
    selected_comments = ' '.join(selected_comments_df["Product Review Comment"].astype(str))
    selected_comments = preprocess_text(selected_comments)
    if not selected_comments.strip():
        continue
    tokenized_words = nltk.word_tokenize(selected_comments)
    stop_words = set(ENGLISH_STOP_WORDS)
    filtered_words = [word for word in tokenized_words if word not in stop_words]
    lemmatizer = nltk.WordNetLemmatizer()
    lemmatized_words = [lemmatizer.lemmatize(word) for word in filtered_words]
    max_phrase_length = 4
    phrases = [phrase for n in range(2, max_phrase_length + 1) for phrase in ngrams(lemmatized_words, n)]
    word_counter = Counter(lemmatized_words)
    phrase_counter = Counter(phrases)

    # Get the top N words and phrases
    top_words = dict(word_counter.most_common(top_n))
    top_phrases = dict(phrase_counter.most_common(top_n))

    # Extract record_id for each Product_ID
    record_ids = selected_comments_df["record_id"].values[0]

    word_counts[(selected_product_id, record_ids)] = top_words
    phrase_counts[(selected_product_id, record_ids)] = top_phrases

word_result_data = []
phrase_result_data = []

for (product_id, record_id), top_words in word_counts.items():
    for word, count in top_words.items():
        word_result_data.append([product_id, record_id, word, count])
for (product_id, record_id), top_phrases in phrase_counts.items():
    for phrase, count in top_phrases.items():
        phrase_result_data.append([product_id, record_id, phrase, count])

word_df = pd.DataFrame(word_result_data, columns=['Product_ID', 'record_id', 'Word', 'Count'])
phrase_df = pd.DataFrame(phrase_result_data, columns=['Product_ID', 'record_id', 'Phrase', 'Count'])

word_df.to_csv('top_words_counts.csv', index=False)
phrase_df.to_csv('top_phrases_counts.csv', index=False)

问题原因
  1. Record_ID关联错误:代码中record_ids = selected_comments_df["record_id"].values[0]仅取每个Product_ID对应的第一条record_id,但一个产品对应多条评论(多个Record_ID),既丢失了关联关系,也会导致统计内容混淆。
  2. 短语统计逻辑错误:将同一产品的所有评论拼接后生成ngrams,会出现跨评论的无意义短语(比如前一条评论的最后一个词和后一条评论的第一个词组成短语),完全不符合实际评论中的短语出现场景。
  3. 词形还原不精准:未指定词性的词形还原默认按名词处理,导致动词、形容词类词汇还原错误,无法统一同一词汇的不同变形。
  4. 统计维度错误:按单个产品取top_n后汇总,而非全局统计后取top_n,导致全局高频的词/短语可能在单个产品中未进入前100,最终结果缺失。

优化后的代码实现
import pandas as pd
from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS
import string
from collections import Counter
from nltk.util import ngrams
import nltk
# 一次性下载所需NLTK资源
nltk.download(['punkt', 'wordnet', 'averaged_perceptron_tagger'])

df = pd.read_excel('productsvydata.xlsx')
# 统一列名,避免大小写/拼写不一致问题
df.rename(columns={'Product Review Comment': 'review_comment', 'Record_ID': 'record_id'}, inplace=True)

def preprocess_text(text):
    if pd.isna(text):
        return ""
    # 转小写
    text = text.lower()
    # 移除标点符号
    translator = str.maketrans('', '', string.punctuation)
    text = text.translate(translator)
    return text

def lemmatize_with_pos(word, pos_tag):
    lemmatizer = nltk.WordNetLemmatizer()
    # 映射NLTK词性到WordNet支持的词性
    pos_map = {'N': 'n', 'V': 'v', 'J': 'a', 'R': 'r'}
    target_pos = pos_map.get(pos_tag[0], 'n')
    return lemmatizer.lemmatize(word, pos=target_pos)

# 全局词频/短语频计数器
global_word_counter = Counter()
global_phrase_counter = Counter()
# 存储每条记录的词/短语统计,用于关联Product_ID和record_id
record_word_map = {}
record_phrase_map = {}

# 逐条处理评论,避免跨评论生成短语
for idx, row in df.iterrows():
    product_id = row['Product_ID']
    record_id = row['record_id']
    comment = row['review_comment']
    
    processed_text = preprocess_text(comment)
    if not processed_text.strip():
        continue
    
    # 分词
    tokens = nltk.word_tokenize(processed_text)
    # 过滤停用词
    filtered_tokens = [token for token in tokens if token not in ENGLISH_STOP_WORDS]
    # 词性标注后精准词形还原
    pos_tags = nltk.pos_tag(filtered_tokens)
    lemmatized_tokens = [lemmatize_with_pos(word, tag) for word, tag in pos_tags]
    
    # 更新全局词频计数器,同时记录当前记录的词频
    word_counts = Counter(lemmatized_tokens)
    global_word_counter.update(word_counts)
    record_word_map[(product_id, record_id)] = word_counts
    
    # 仅在当前评论内生成2-4元短语
    max_phrase_len = 4
    phrases = []
    for n in range(2, max_phrase_len + 1):
        phrases.extend(ngrams(lemmatized_tokens, n))
    phrase_counts = Counter(phrases)
    global_phrase_counter.update(phrase_counts)
    record_phrase_map[(product_id, record_id)] = phrase_counts

# 获取全局top_n的词和短语
top_n = 100
top_global_words = set([word for word, _ in global_word_counter.most_common(top_n)])
top_global_phrases = set([phrase for phrase, _ in global_phrase_counter.most_common(top_n)])

# 构建结果数据,仅保留全局top_n的内容,同时保留关联关系
word_result_data = []
for (product_id, record_id), word_counts in record_word_map.items():
    for word, count in word_counts.items():
        if word in top_global_words:
            word_result_data.append([product_id, record_id, word, count])

phrase_result_data = []
for (product_id, record_id), phrase_counts in record_phrase_map.items():
    for phrase, count in phrase_counts.items():
        if phrase in top_global_phrases:
            phrase_result_data.append([product_id, record_id, phrase, count])

# 转换为DataFrame并导出
word_df = pd.DataFrame(word_result_data, columns=['Product_ID', 'record_id', 'Word', 'Count'])
phrase_df = pd.DataFrame(phrase_result_data, columns=['Product_ID', 'record_id', 'Phrase', 'Count'])

# 按产品+记录+词/短语分组汇总,确保计数准确
word_df_grouped = word_df.groupby(['Product_ID', 'record_id', 'Word'])['Count'].sum().reset_index()
phrase_df_grouped = phrase_df.groupby(['Product_ID', 'record_id', 'Phrase'])['Count'].sum().reset_index()

word_df_grouped.to_csv('top_words_counts.csv', index=False)
phrase_df_grouped.to_csv('top_phrases_counts.csv', index=False)

优化说明
  1. 修复关联关系:逐条处理评论,完整保留每个Product_ID与record_id的对应关系,不会丢失记录关联。
  2. 提升预处理精度:添加词性标注后再进行词形还原,确保不同词性的词汇被正确统一(比如动词service和名词service不会被错误还原)。
  3. 修正短语统计:仅在单条评论内生成ngrams,避免跨评论生成无意义短语,保证短语统计的真实性。
  4. 优化统计逻辑:先全局统计所有词和短语的频次,再筛选top_n,确保全局高频内容不会被单个产品的top_n限制过滤。
  5. 性能优化:逐条处理减少内存占用,全局统计后过滤top_n,在保证准确性的同时控制运行效率。

内容的提问来源于stack exchange,提问作者user14452102

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 03:54:52