You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rake-nltk提取关键词后,如何定位结果对应的原始数据行?

How to Track Rake-NLTK Keyword Phrases Back to Original Data Rows

Great question! The default extract_keywords_from_sentences method in Rake-NLTK processes all input sentences in bulk, which means it doesn’t retain any information about which keyword phrases come from which original row in your DataFrame. To fix this, you’ll need to process each sentence individually while keeping track of its index.

Here’s a modified version of your code that preserves the original row association:

from rake_nltk import Rake
import pandas as pd

# Initialize your Rake object with your desired parameters
r = Rake(include_repeated_phrases=False, min_length=2, max_length=3)

# List to store results: each entry is (score, keyword_phrase, original_row_index)
ranked_phrases_with_row_data = []

# Iterate over each row in your DataFrame, keeping track of the index
for idx, sentence in enumerate(data['five_gram']):
    # Skip empty entries to avoid errors
    if pd.isna(sentence):
        continue
    
    # Reset the Rake object to clear any leftover data from previous iterations
    r.reset()
    # Extract keywords from the current sentence
    r.extract_keywords_from_text(sentence)
    # Get ranked phrases for this sentence and filter by score
    for score, phrase in r.get_ranked_phrases_with_scores():
        if score > 5:
            ranked_phrases_with_row_data.append((score, phrase, idx))

# Print the results—each entry now includes the original row index
print(ranked_phrases_with_row_data)

Key Details:

  • We use enumerate() to get both the row index (idx) and the sentence text from your five_gram column.
  • The r.reset() call ensures the Rake object doesn’t carry over data from previous sentences.
  • We skip empty values with pd.isna(sentence) to prevent errors when processing missing text.

If You’re Using Custom DataFrame Indices:

If your DataFrame uses non-default indices (e.g., strings or non-sequential integers), use iterrows() instead to capture the actual row index:

for actual_idx, row in data.iterrows():
    sentence = row['five_gram']
    if pd.notna(sentence):
        r.reset()
        r.extract_keywords_from_text(sentence)
        for score, phrase in r.get_ranked_phrases_with_scores():
            if score > 5:
                ranked_phrases_with_row_data.append((score, phrase, actual_idx))

Once you have the index, you can easily retrieve the original row from your DataFrame using data.iloc[idx] (for default integer indices) or data.loc[actual_idx] (for custom indices).

内容的提问来源于stack exchange,提问作者Black_Stallion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:48:10