You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中将列中指定词汇替换为列表内所有词汇并生成新列

Solution for Generating New Text Column with Number Replacements

Let's solve your problem step by step. The goal is to create a new_text column where each sentence containing words from my_list gets expanded into all possible variation sentences (swapping matched words with every entry in my_list), while sentences without target words stay unchanged.

Step 1: Setup and Define Data

First, let's prepare our libraries and original data:

import pandas as pd
import re

# Your original dataset
data = {"text":["I have one apple and two bananas", "this is my apple", "she has three apples","My friend has five apples but she only has one banana"]}
df = pd.DataFrame(data=data, columns=['text'])
my_list = ['one','two','three','four','five']

Step 2: Create a Custom Processing Function

We'll build a function to handle each sentence individually:

  1. Identifies all target words from my_list in the sentence (matching whole words, case-insensitively)
  2. Generates every possible replacement sentence by swapping each matched word with every entry in my_list
  3. Returns the original sentence if no matches are found, otherwise joins all generated sentences into a single string
def generate_new_text(sentence, num_list):
    # Regex pattern to match whole words from num_list, case-insensitive
    num_pattern = re.compile(r'\b(' + '|'.join(num_list) + r')\b', re.IGNORECASE)
    # Find all matching numbers in the current sentence
    matched_nums = num_pattern.findall(sentence)
    
    if not matched_nums:
        # No target words found? Return the original sentence
        return sentence
    
    # Collect unique generated sentences (use a set to avoid duplicates)
    generated_sentences = set()
    for match in matched_nums:
        for num in num_list:
            # Replace one occurrence of the matched number at a time
            new_sentence = num_pattern.sub(num, sentence, count=1)
            generated_sentences.add(new_sentence)
    
    # Join all unique sentences into a comma-separated string
    return ', '.join(generated_sentences)

Step 3: Apply the Function to Your DataFrame

Now we'll apply the function to the text column to create the new_text column:

df['new_text'] = df['text'].apply(lambda x: generate_new_text(x, my_list))

Step 4: Verify the Result

When you print the DataFrame:

print(df)

You'll get output similar to your expected example (truncated for readability):

text                                                                 new_text
0                I have one apple and two bananas  I have two apple and two bananas, I have five apple and two bananas, I...
1                                this is my apple                                                                this is my apple
2                            she has three apples  she has four apples, she has five apples, she has two apples, she ha...
3  My friend has five apples but she only has one banana  My friend has five apples but she only has two banana, My friend...

Key Details:

  • Whole Word Matching: The regex uses \b to ensure we only match full words (so "ones" won't be incorrectly matched to "one")
  • Case Insensitivity: re.IGNORECASE handles matches like "One" or "ONE" in the original text
  • Duplicate Prevention: Using a set removes duplicate sentences (you can omit this if duplicates are acceptable per your note)
  • Full List Coverage: By looping through every entry in my_list for each matched word, we guarantee every number in your list appears in the generated sentences

内容的提问来源于stack exchange,提问作者lulu mirzai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 08:27:48