在Pandas中按Product_id分组,根据退款状态生成含统一产品名称的随机评论
Got it, let's fix this issue where the same product_id ends up with different product names in reviews. Here's a straightforward approach that keeps your random comment logic intact but ensures consistency per product:
Step 1: Create a fixed product name mapping for each product_id
First, we'll generate a dictionary that maps every unique product_id to a single, randomly selected product name. This way, all reviews for the same product will use the same name:
import random # Assume your existing products list is defined somewhere products = ["rocking char", "coffee table", "floor lamp"] # example list # Create a mapping: each product_id gets one fixed product name product_name_map = {pid: random.choice(products) for pid in df['product_id'].unique()}
Step 2: Update your review functions to accept a fixed product name
Modify your random_negative_sentence and random_positive_sentence functions to take a product_name parameter instead of randomly picking one internally. This lets us pass in the fixed name from our mapping:
Updated Negative Review Function
def random_negative_sentence(product_name): # Keep your existing random lists (negative, colors, person, etc.) adj = random.choice(negative) color = random.choice(colors) who = random.choice(person) verb = random.choice(negative_verbs) end = random.choice(l_neg) sentence1 = f'{who} {verb} the {color} {adj} {product_name}. {end}.' sentence2 = f'The {color} {product_name} is so {adj} that even {who} {verb} it! {end}' sentence3 = f'{who} {verb} the {color} {product_name} because it is so {adj}! {end}' sentence4 = f'{who} {verb} the {color} {product_name} because it is {adj} and {random.choice(negative)}! {end}' return random.choice([sentence1, sentence2, sentence3, sentence4])
Updated Positive Review Function
def random_positive_sentence(product_name): # Use your existing positive word lists adj = random.choice(positive_adjectives) color = random.choice(colors) who = random.choice(person) verb = random.choice(positive_verbs) end = random.choice(l_pos) sentence1 = f'{who} {verb} the {color} {adj} {product_name}. {end}.' sentence2 = f'The {color} {product_name} is so {adj} that even {who} {verb} it! {end}' sentence3 = f'{who} {verb} the {color} {product_name} because it is so {adj}! {end}' sentence4 = f'{who} {verb} the {color} {product_name} because it is {adj} and {random.choice(positive_adjectives)}! {end}' return random.choice([sentence1, sentence2, sentence3, sentence4])
Step 3: Generate reviews using the fixed product names
Now use df.apply() to generate reviews row by row. For each row, we pull the fixed product name from our mapping, then call the appropriate positive/negative function based on the refund status:
df['review'] = df.apply( lambda row: random_positive_sentence(product_name_map[row['product_id']]) if not row['refund'] else random_negative_sentence(product_name_map[row['product_id']]), axis=1 )
How This Works
- The
product_name_mapensures everyproduct_idhas one consistent product name across all its reviews. - Your original randomization logic for adjectives, colors, people, verbs, and sentence structures is preserved—so reviews for the same product will still be unique, just with the same product name.
- The
applycall checks therefundstatus for each row and uses the correct review function with the fixed product name.
Testing this with your sample DataFrame will result in all rows with product_id = WFL2ILKU3Z using the same product name (like "rocking char") in their reviews, while the rest of the comment content remains random and aligned with the refund status.
内容的提问来源于stack exchange,提问作者Jonas Palačionis

