如何用另一列内容替换句子任意位置的指定字符串
Solution for Replacing Strings Anywhere in Sentences
Got it, let's get this sorted so you can replace matching terms anywhere in your sentences, not just the first word. Here's a straightforward approach that works with your existing NLTK setup:
Step-by-Step Explanation
- Tokenize the Sentence: Split your input sentence into individual words using NLTK's
word_tokenize—this lets us check every word, no matter its position. - Case-Insensitive Matching: Convert each token to lowercase to compare against your all-lowercase
col1values (so "Aaa1" matches "aaa1"). - Replace Matching Tokens: Iterate through each token, replacing it with your target value (like "NNP") if its lowercase version exists in
col1. - Reconstruct the Sentence: Join the processed tokens back into a single string.
Code Implementation
from nltk import word_tokenize def replace_any_position(sentence, target_lower_terms, replacement): # Split sentence into individual words tokens = word_tokenize(sentence) # Check each token: replace if lowercase matches any term in target list processed_tokens = [ replacement if token.lower() in target_lower_terms else token for token in tokens ] # Join tokens back into a sentence return ' '.join(processed_tokens) # Example usage if __name__ == "__main__": # Your col1 values (all lowercase) col1 = ["aaa1"] # Test case 1: Target in middle of sentence input_sentence1 = "Hello Aaa1 my very good friend" output1 = replace_any_position(input_sentence1, col1, "NNP") print(output1) # Output: Hello NNP my very good friend # Test case 2: Target at start of sentence input_sentence2 = "Aaa1 my very good friend" output2 = replace_any_position(input_sentence2, col1, "NNP") print(output2) # Output: NNP my very good friend
Notes for DataFrame Integration
If your col1 is a column in a pandas DataFrame, convert it to a set for faster lookups (sets have O(1) membership checks vs lists' O(n)):
import pandas as pd # Example DataFrame df = pd.DataFrame({ "col1": ["aaa1", "bbb2"], "col2": ["Hello Aaa1 my friend", "Bbb2 is here"] }) # Convert col1 to a set of lowercase terms target_set = set(df["col1"].str.lower()) # Apply the function to each row in col2 df["col2_processed"] = df["col2"].apply( lambda x: replace_any_position(x, target_set, "NNP") ) print(df["col2_processed"])
This will handle replacements anywhere in the sentence, exactly what you need!
内容的提问来源于stack exchange,提问作者belriso
相关产品推荐
相关产品推荐

