Python词频统计结果失真,求问题排查与优化方案
问题
作为NLP初学者,我有一份超20000行的产品评论数据表(示例如下),已编写Python代码实现评论的词频及短语频次统计并导出CSV。为避免运行卡顿设置了top_n=100,但统计结果严重失真:比如原数据中‘customer service’提及次数远多于统计结果,‘customer’计数仅为9次。需要排查问题原因并提供代码优化方案,同时要关联Product_ID与record_id适配工作工具。
产品评论表示例
| Record_ID | Product_ID | Review Comment |
|---|---|---|
| 1234 | 89847457 | I love this product it was shipped fast and is comfortable |
当前代码实现
import pandas as pd from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS import string from collections import Counter from nltk.util import ngrams import nltk nltk.download('punkt') df = pd.read_excel('productsvydata.xlsx') def preprocess_text(text): translator = str.maketrans('', '', string.punctuation) text = text.lower() text = text.translate(translator) return text word_counts = {} phrase_counts = {} unique_product_ids = df["Product_ID"].unique() # Set the number of top words and phrases you want to keep top_n = 100 for selected_product_id in unique_product_ids: selected_comments_df = df[df["Product_ID"] == selected_product_id] selected_comments = ' '.join(selected_comments_df["Product Review Comment"].astype(str)) selected_comments = preprocess_text(selected_comments) if not selected_comments.strip(): continue tokenized_words = nltk.word_tokenize(selected_comments) stop_words = set(ENGLISH_STOP_WORDS) filtered_words = [word for word in tokenized_words if word not in stop_words] lemmatizer = nltk.WordNetLemmatizer() lemmatized_words = [lemmatizer.lemmatize(word) for word in filtered_words] max_phrase_length = 4 phrases = [phrase for n in range(2, max_phrase_length + 1) for phrase in ngrams(lemmatized_words, n)] word_counter = Counter(lemmatized_words) phrase_counter = Counter(phrases) # Get the top N words and phrases top_words = dict(word_counter.most_common(top_n)) top_phrases = dict(phrase_counter.most_common(top_n)) # Extract record_id for each Product_ID record_ids = selected_comments_df["record_id"].values[0] word_counts[(selected_product_id, record_ids)] = top_words phrase_counts[(selected_product_id, record_ids)] = top_phrases word_result_data = [] phrase_result_data = [] for (product_id, record_id), top_words in word_counts.items(): for word, count in top_words.items(): word_result_data.append([product_id, record_id, word, count]) for (product_id, record_id), top_phrases in phrase_counts.items(): for phrase, count in top_phrases.items(): phrase_result_data.append([product_id, record_id, phrase, count]) word_df = pd.DataFrame(word_result_data, columns=['Product_ID', 'record_id', 'Word', 'Count']) phrase_df = pd.DataFrame(phrase_result_data, columns=['Product_ID', 'record_id', 'Phrase', 'Count']) word_df.to_csv('top_words_counts.csv', index=False) phrase_df.to_csv('top_phrases_counts.csv', index=False)
问题原因
- Record_ID关联错误:代码中
record_ids = selected_comments_df["record_id"].values[0]仅取每个Product_ID对应的第一条record_id,但一个产品对应多条评论(多个Record_ID),既丢失了关联关系,也会导致统计内容混淆。 - 短语统计逻辑错误:将同一产品的所有评论拼接后生成ngrams,会出现跨评论的无意义短语(比如前一条评论的最后一个词和后一条评论的第一个词组成短语),完全不符合实际评论中的短语出现场景。
- 词形还原不精准:未指定词性的词形还原默认按名词处理,导致动词、形容词类词汇还原错误,无法统一同一词汇的不同变形。
- 统计维度错误:按单个产品取top_n后汇总,而非全局统计后取top_n,导致全局高频的词/短语可能在单个产品中未进入前100,最终结果缺失。
优化后的代码实现
import pandas as pd from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS import string from collections import Counter from nltk.util import ngrams import nltk # 一次性下载所需NLTK资源 nltk.download(['punkt', 'wordnet', 'averaged_perceptron_tagger']) df = pd.read_excel('productsvydata.xlsx') # 统一列名,避免大小写/拼写不一致问题 df.rename(columns={'Product Review Comment': 'review_comment', 'Record_ID': 'record_id'}, inplace=True) def preprocess_text(text): if pd.isna(text): return "" # 转小写 text = text.lower() # 移除标点符号 translator = str.maketrans('', '', string.punctuation) text = text.translate(translator) return text def lemmatize_with_pos(word, pos_tag): lemmatizer = nltk.WordNetLemmatizer() # 映射NLTK词性到WordNet支持的词性 pos_map = {'N': 'n', 'V': 'v', 'J': 'a', 'R': 'r'} target_pos = pos_map.get(pos_tag[0], 'n') return lemmatizer.lemmatize(word, pos=target_pos) # 全局词频/短语频计数器 global_word_counter = Counter() global_phrase_counter = Counter() # 存储每条记录的词/短语统计,用于关联Product_ID和record_id record_word_map = {} record_phrase_map = {} # 逐条处理评论,避免跨评论生成短语 for idx, row in df.iterrows(): product_id = row['Product_ID'] record_id = row['record_id'] comment = row['review_comment'] processed_text = preprocess_text(comment) if not processed_text.strip(): continue # 分词 tokens = nltk.word_tokenize(processed_text) # 过滤停用词 filtered_tokens = [token for token in tokens if token not in ENGLISH_STOP_WORDS] # 词性标注后精准词形还原 pos_tags = nltk.pos_tag(filtered_tokens) lemmatized_tokens = [lemmatize_with_pos(word, tag) for word, tag in pos_tags] # 更新全局词频计数器,同时记录当前记录的词频 word_counts = Counter(lemmatized_tokens) global_word_counter.update(word_counts) record_word_map[(product_id, record_id)] = word_counts # 仅在当前评论内生成2-4元短语 max_phrase_len = 4 phrases = [] for n in range(2, max_phrase_len + 1): phrases.extend(ngrams(lemmatized_tokens, n)) phrase_counts = Counter(phrases) global_phrase_counter.update(phrase_counts) record_phrase_map[(product_id, record_id)] = phrase_counts # 获取全局top_n的词和短语 top_n = 100 top_global_words = set([word for word, _ in global_word_counter.most_common(top_n)]) top_global_phrases = set([phrase for phrase, _ in global_phrase_counter.most_common(top_n)]) # 构建结果数据,仅保留全局top_n的内容,同时保留关联关系 word_result_data = [] for (product_id, record_id), word_counts in record_word_map.items(): for word, count in word_counts.items(): if word in top_global_words: word_result_data.append([product_id, record_id, word, count]) phrase_result_data = [] for (product_id, record_id), phrase_counts in record_phrase_map.items(): for phrase, count in phrase_counts.items(): if phrase in top_global_phrases: phrase_result_data.append([product_id, record_id, phrase, count]) # 转换为DataFrame并导出 word_df = pd.DataFrame(word_result_data, columns=['Product_ID', 'record_id', 'Word', 'Count']) phrase_df = pd.DataFrame(phrase_result_data, columns=['Product_ID', 'record_id', 'Phrase', 'Count']) # 按产品+记录+词/短语分组汇总,确保计数准确 word_df_grouped = word_df.groupby(['Product_ID', 'record_id', 'Word'])['Count'].sum().reset_index() phrase_df_grouped = phrase_df.groupby(['Product_ID', 'record_id', 'Phrase'])['Count'].sum().reset_index() word_df_grouped.to_csv('top_words_counts.csv', index=False) phrase_df_grouped.to_csv('top_phrases_counts.csv', index=False)
优化说明
- 修复关联关系:逐条处理评论,完整保留每个
Product_ID与record_id的对应关系,不会丢失记录关联。 - 提升预处理精度:添加词性标注后再进行词形还原,确保不同词性的词汇被正确统一(比如动词
service和名词service不会被错误还原)。 - 修正短语统计:仅在单条评论内生成ngrams,避免跨评论生成无意义短语,保证短语统计的真实性。
- 优化统计逻辑:先全局统计所有词和短语的频次,再筛选top_n,确保全局高频内容不会被单个产品的top_n限制过滤。
- 性能优化:逐条处理减少内存占用,全局统计后过滤top_n,在保证准确性的同时控制运行效率。
内容的提问来源于stack exchange,提问作者user14452102
相关产品推荐
相关产品推荐

