使用rake-nltk提取关键词后,如何定位结果对应的原始数据行?
Great question! The default extract_keywords_from_sentences method in Rake-NLTK processes all input sentences in bulk, which means it doesn’t retain any information about which keyword phrases come from which original row in your DataFrame. To fix this, you’ll need to process each sentence individually while keeping track of its index.
Here’s a modified version of your code that preserves the original row association:
from rake_nltk import Rake import pandas as pd # Initialize your Rake object with your desired parameters r = Rake(include_repeated_phrases=False, min_length=2, max_length=3) # List to store results: each entry is (score, keyword_phrase, original_row_index) ranked_phrases_with_row_data = [] # Iterate over each row in your DataFrame, keeping track of the index for idx, sentence in enumerate(data['five_gram']): # Skip empty entries to avoid errors if pd.isna(sentence): continue # Reset the Rake object to clear any leftover data from previous iterations r.reset() # Extract keywords from the current sentence r.extract_keywords_from_text(sentence) # Get ranked phrases for this sentence and filter by score for score, phrase in r.get_ranked_phrases_with_scores(): if score > 5: ranked_phrases_with_row_data.append((score, phrase, idx)) # Print the results—each entry now includes the original row index print(ranked_phrases_with_row_data)
Key Details:
- We use
enumerate()to get both the row index (idx) and the sentence text from yourfive_gramcolumn. - The
r.reset()call ensures the Rake object doesn’t carry over data from previous sentences. - We skip empty values with
pd.isna(sentence)to prevent errors when processing missing text.
If You’re Using Custom DataFrame Indices:
If your DataFrame uses non-default indices (e.g., strings or non-sequential integers), use iterrows() instead to capture the actual row index:
for actual_idx, row in data.iterrows(): sentence = row['five_gram'] if pd.notna(sentence): r.reset() r.extract_keywords_from_text(sentence) for score, phrase in r.get_ranked_phrases_with_scores(): if score > 5: ranked_phrases_with_row_data.append((score, phrase, actual_idx))
Once you have the index, you can easily retrieve the original row from your DataFrame using data.iloc[idx] (for default integer indices) or data.loc[actual_idx] (for custom indices).
内容的提问来源于stack exchange,提问作者Black_Stallion

