You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取DataFrame中总计30个高频n-grams转为列并填充TF-IDF分数?

解决步骤

核心思路

先统一收集所有单/双/三元组,统计出总计30个最频繁的n-grams,再针对这些目标n-gram计算每篇文章的TF-IDF分数,最终合并到原DataFrame。


步骤1:导入依赖库

import pandas as pd
from collections import Counter
from sklearn.feature_extraction.text import TfidfVectorizer
# 确保已导入nltk相关模块并下载所需资源
import nltk
from nltk.tokenize import word_tokenize
from nltk.stem import WordNetLemmatizer
from nltk.corpus import stopwords
nltk.download('punkt')
nltk.download('wordnet')
nltk.download('stopwords')
stop_words = set(stopwords.words('english'))

步骤2:统一处理n-gram格式(元组转字符串)

双/三元组是元组格式,无法直接作为列名或统计,先转成下划线连接的字符串:

def ngram_to_str(ngram):
    # 处理双/三元组元组
    if isinstance(ngram, tuple):
        return "_".join(ngram)
    # 处理单字组字符串
    return ngram

步骤3:统计所有n-gram的频率,筛选Top30

遍历所有文章收集所有n-gram,用Counter统计频率后取前30:

all_ngrams = []
for _, row in press.iterrows():
    # 收集单字组
    all_ngrams.extend([ngram_to_str(ug) for ug in row['unigrams']])
    # 收集双字组
    all_ngrams.extend([ngram_to_str(bg) for bg in row['bigrams']])
    # 收集三字组
    all_ngrams.extend([ngram_to_str(tg) for tg in row['trigrams']])

# 统计频率并取Top30
ngram_counter = Counter(all_ngrams)
top_30_ngrams = [ngram for ngram, _ in ngram_counter.most_common(30)]

步骤4:生成每篇文章的n-gram文本表示

把每篇文章的所有n-gram转成空格分隔的字符串,适配TF-IDF工具的输入要求:

def article_to_ngram_str(row):
    ngram_list = []
    ngram_list.extend([ngram_to_str(ug) for ug in row['unigrams']])
    ngram_list.extend([ngram_to_str(bg) for bg in row['bigrams']])
    ngram_list.extend([ngram_to_str(tg) for tg in row['trigrams']])
    return ' '.join(ngram_list)

# 新增临时列存储所有n-gram的字符串拼接
press['all_ngrams_text'] = press.apply(article_to_ngram_str, axis=1)

步骤5:计算TF-IDF并合并到原DataFrame

指定TF-IDF工具只计算Top30 n-gram的分数,然后把结果合并到原表:

# 初始化TF-IDF向量器,限定词汇为Top30 n-gram
tfidf = TfidfVectorizer(vocabulary=top_30_ngrams)
# 计算TF-IDF矩阵
tfidf_matrix = tfidf.fit_transform(press['all_ngrams_text'])
# 转成DataFrame格式
tfidf_df = pd.DataFrame(
    tfidf_matrix.toarray(),
    columns=tfidf.get_feature_names_out(),
    index=press.index
)
# 合并到原DataFrame
press = pd.concat([press, tfidf_df], axis=1)
# 删除临时列
press.drop('all_ngrams_text', axis=1, inplace=True)

内容的提问来源于stack exchange,提问作者Dave

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 08:15:44