如何根据@符号后方单词将其替换为基于DataFrame中位数与标准差生成的数值
Let’s walk through a practical, Python-based solution using pandas and numpy to tackle this problem. Here’s how you can make this work smoothly:
Step 1: Extract the Target Word After @
First, we need to pull out the word that follows the @ symbol in your generated sentence. Regular expressions are perfect here—they handle cases where there might be spaces between @ and the word.
import re generated_sentence = "Had to wait @ minutes for the pizza" # Match @ followed by optional whitespace, then capture the word match = re.search(r'@\s*(\w+)', generated_sentence) if match: target_word = match.group(1) # This will be "minutes" in your example else: # Handle cases where @ isn't found or no word follows it target_word = None print("No valid @ + word pattern found in the sentence")
Step 2: Fetch Median and Std from Your DataFrame
Assuming your DataFrame is named stats_df, we’ll look up the corresponding stats for the target word. We’ll add a check in case the word isn’t present in the DataFrame to avoid errors.
import pandas as pd # Example DataFrame (replace with your actual dataset) stats_df = pd.DataFrame({ 'word': ['minutes', 'pm', 'stars'], 'median': [20, 9, 3], 'std': [20, 2, 2] }) if target_word and target_word in stats_df['word'].values: # Get the row for the target word word_stats = stats_df.loc[stats_df['word'] == target_word].iloc[0] median_val = word_stats['median'] std_val = word_stats['std'] else: # Fallback values if the word isn't in your DataFrame median_val = 10 std_val = 5 print(f"Word '{target_word}' not found in DataFrame—using fallback stats")
Step 3: Generate a Valid Numeric Value
We’ll generate a value from a normal distribution using the median (as the mean) and std from your DataFrame. Since we’re dealing with real-world values like wait times or ratings, we’ll ensure the result is a non-negative integer.
import numpy as np # Generate a normally distributed value, round to integer, ensure it's non-negative generated_num = int(np.round(np.random.normal(loc=median_val, scale=std_val))) generated_num = max(0, generated_num) # Avoid nonsensical negative numbers
Step 4: Replace @ with the Generated Number
Finally, swap out the @ (and any surrounding whitespace) with our generated number to fix the sentence.
# Replace @ and optional whitespace with the generated number corrected_sentence = re.sub(r'@\s*', f'{generated_num} ', generated_sentence) print(corrected_sentence) # Example output: "Had to wait 18 minutes for the pizza"
Reusable Function for Multiple Sentences
To make this scalable, wrap everything in a function that takes the sentence and DataFrame as inputs:
import re import pandas as pd import numpy as np def replace_at_with_generated_value(sentence, stats_df, fallback_median=10, fallback_std=5): # Extract target word match = re.search(r'@\s*(\w+)', sentence) if not match: return sentence # Return original if no @ pattern found target_word = match.group(1) # Get stats from DataFrame if target_word in stats_df['word'].values: word_stats = stats_df.loc[stats_df['word'] == target_word].iloc[0] median_val = word_stats['median'] std_val = word_stats['std'] else: median_val = fallback_median std_val = fallback_std # Generate valid number generated_num = int(np.round(np.random.normal(loc=median_val, scale=std_val))) generated_num = max(0, generated_num) # Replace @ in sentence corrected_sentence = re.sub(r'@\s*', f'{generated_num} ', sentence) return corrected_sentence # Test it out with your example stats_df = pd.DataFrame({ 'word': ['minutes', 'pm', 'stars'], 'median': [20, 9, 3], 'std': [20, 2, 2] }) test_sentence = "Had to wait @ minutes for the pizza" print(replace_at_with_generated_value(test_sentence, stats_df))
Quick Notes on Edge Cases
- Negative values: Using
max(0, generated_num)ensures we don’t get nonsensical values like "-5 minutes". - Missing words: The fallback stats let your code keep running even if a word isn’t in your DataFrame—adjust these defaults to fit your use case.
- Multiple @ symbols: If your sentence has multiple
@s, this function will replace all of them with the same value based on the first matched word. To handle multiple unique@ + wordpairs, modify the regex to find all matches and process each one individually.
内容的提问来源于stack exchange,提问作者FlubberBeer

