如何在Pandas中将列中指定词汇替换为列表内所有词汇并生成新列
Solution for Generating New Text Column with Number Replacements
Let's solve your problem step by step. The goal is to create a new_text column where each sentence containing words from my_list gets expanded into all possible variation sentences (swapping matched words with every entry in my_list), while sentences without target words stay unchanged.
Step 1: Setup and Define Data
First, let's prepare our libraries and original data:
import pandas as pd import re # Your original dataset data = {"text":["I have one apple and two bananas", "this is my apple", "she has three apples","My friend has five apples but she only has one banana"]} df = pd.DataFrame(data=data, columns=['text']) my_list = ['one','two','three','four','five']
Step 2: Create a Custom Processing Function
We'll build a function to handle each sentence individually:
- Identifies all target words from
my_listin the sentence (matching whole words, case-insensitively) - Generates every possible replacement sentence by swapping each matched word with every entry in
my_list - Returns the original sentence if no matches are found, otherwise joins all generated sentences into a single string
def generate_new_text(sentence, num_list): # Regex pattern to match whole words from num_list, case-insensitive num_pattern = re.compile(r'\b(' + '|'.join(num_list) + r')\b', re.IGNORECASE) # Find all matching numbers in the current sentence matched_nums = num_pattern.findall(sentence) if not matched_nums: # No target words found? Return the original sentence return sentence # Collect unique generated sentences (use a set to avoid duplicates) generated_sentences = set() for match in matched_nums: for num in num_list: # Replace one occurrence of the matched number at a time new_sentence = num_pattern.sub(num, sentence, count=1) generated_sentences.add(new_sentence) # Join all unique sentences into a comma-separated string return ', '.join(generated_sentences)
Step 3: Apply the Function to Your DataFrame
Now we'll apply the function to the text column to create the new_text column:
df['new_text'] = df['text'].apply(lambda x: generate_new_text(x, my_list))
Step 4: Verify the Result
When you print the DataFrame:
print(df)
You'll get output similar to your expected example (truncated for readability):
text new_text 0 I have one apple and two bananas I have two apple and two bananas, I have five apple and two bananas, I... 1 this is my apple this is my apple 2 she has three apples she has four apples, she has five apples, she has two apples, she ha... 3 My friend has five apples but she only has one banana My friend has five apples but she only has two banana, My friend...
Key Details:
- Whole Word Matching: The regex uses
\bto ensure we only match full words (so "ones" won't be incorrectly matched to "one") - Case Insensitivity:
re.IGNORECASEhandles matches like "One" or "ONE" in the original text - Duplicate Prevention: Using a
setremoves duplicate sentences (you can omit this if duplicates are acceptable per your note) - Full List Coverage: By looping through every entry in
my_listfor each matched word, we guarantee every number in your list appears in the generated sentences
内容的提问来源于stack exchange,提问作者lulu mirzai
相关产品推荐
相关产品推荐

