You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何根据@符号后方单词将其替换为基于DataFrame中位数与标准差生成的数值

Solution to Replace @ with Generated Value Based on DataFrame Stats

Let’s walk through a practical, Python-based solution using pandas and numpy to tackle this problem. Here’s how you can make this work smoothly:

Step 1: Extract the Target Word After @

First, we need to pull out the word that follows the @ symbol in your generated sentence. Regular expressions are perfect here—they handle cases where there might be spaces between @ and the word.

import re

generated_sentence = "Had to wait @ minutes for the pizza"

# Match @ followed by optional whitespace, then capture the word
match = re.search(r'@\s*(\w+)', generated_sentence)
if match:
    target_word = match.group(1)  # This will be "minutes" in your example
else:
    # Handle cases where @ isn't found or no word follows it
    target_word = None
    print("No valid @ + word pattern found in the sentence")

Step 2: Fetch Median and Std from Your DataFrame

Assuming your DataFrame is named stats_df, we’ll look up the corresponding stats for the target word. We’ll add a check in case the word isn’t present in the DataFrame to avoid errors.

import pandas as pd

# Example DataFrame (replace with your actual dataset)
stats_df = pd.DataFrame({
    'word': ['minutes', 'pm', 'stars'],
    'median': [20, 9, 3],
    'std': [20, 2, 2]
})

if target_word and target_word in stats_df['word'].values:
    # Get the row for the target word
    word_stats = stats_df.loc[stats_df['word'] == target_word].iloc[0]
    median_val = word_stats['median']
    std_val = word_stats['std']
else:
    # Fallback values if the word isn't in your DataFrame
    median_val = 10
    std_val = 5
    print(f"Word '{target_word}' not found in DataFrame—using fallback stats")

Step 3: Generate a Valid Numeric Value

We’ll generate a value from a normal distribution using the median (as the mean) and std from your DataFrame. Since we’re dealing with real-world values like wait times or ratings, we’ll ensure the result is a non-negative integer.

import numpy as np

# Generate a normally distributed value, round to integer, ensure it's non-negative
generated_num = int(np.round(np.random.normal(loc=median_val, scale=std_val)))
generated_num = max(0, generated_num)  # Avoid nonsensical negative numbers

Step 4: Replace @ with the Generated Number

Finally, swap out the @ (and any surrounding whitespace) with our generated number to fix the sentence.

# Replace @ and optional whitespace with the generated number
corrected_sentence = re.sub(r'@\s*', f'{generated_num} ', generated_sentence)
print(corrected_sentence)
# Example output: "Had to wait 18 minutes for the pizza"

Reusable Function for Multiple Sentences

To make this scalable, wrap everything in a function that takes the sentence and DataFrame as inputs:

import re
import pandas as pd
import numpy as np

def replace_at_with_generated_value(sentence, stats_df, fallback_median=10, fallback_std=5):
    # Extract target word
    match = re.search(r'@\s*(\w+)', sentence)
    if not match:
        return sentence  # Return original if no @ pattern found
    
    target_word = match.group(1)
    
    # Get stats from DataFrame
    if target_word in stats_df['word'].values:
        word_stats = stats_df.loc[stats_df['word'] == target_word].iloc[0]
        median_val = word_stats['median']
        std_val = word_stats['std']
    else:
        median_val = fallback_median
        std_val = fallback_std
    
    # Generate valid number
    generated_num = int(np.round(np.random.normal(loc=median_val, scale=std_val)))
    generated_num = max(0, generated_num)
    
    # Replace @ in sentence
    corrected_sentence = re.sub(r'@\s*', f'{generated_num} ', sentence)
    return corrected_sentence

# Test it out with your example
stats_df = pd.DataFrame({
    'word': ['minutes', 'pm', 'stars'],
    'median': [20, 9, 3],
    'std': [20, 2, 2]
})

test_sentence = "Had to wait @ minutes for the pizza"
print(replace_at_with_generated_value(test_sentence, stats_df))

Quick Notes on Edge Cases

  • Negative values: Using max(0, generated_num) ensures we don’t get nonsensical values like "-5 minutes".
  • Missing words: The fallback stats let your code keep running even if a word isn’t in your DataFrame—adjust these defaults to fit your use case.
  • Multiple @ symbols: If your sentence has multiple @s, this function will replace all of them with the same value based on the first matched word. To handle multiple unique @ + word pairs, modify the regex to find all matches and process each one individually.

内容的提问来源于stack exchange,提问作者FlubberBeer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 23:07:44