基于Pandas DataFrame匹配关键词并提取其前至首个逗号的文本
Got it, let's walk through how to solve this problem. You've got a Pandas DataFrame with id, product_details, and taxo columns—where product_details holds comma-separated key-value pairs wrapped in square brackets. Your goal is to find the segment containing a specific keyword, then pull the text from the start of that segment up to the first comma (or the whole segment if there's no comma in it).
Step 1: Set Up Sample Data (to replicate your scenario)
First, let's create a sample DataFrame matching your example to test our solution:
import pandas as pd # Recreate your sample DataFrame data = { 'id': [100, 200], 'product_details': [ "[Sales Package=6 Pair slipper, Strap Material=Rubber, qty=1, categoryPath=Footwear>Men>Slippers & Flip Flops, codAvailable=true, detailedSpecs=Multicolor Color, None Character Type Slippers For Men Sole Material Rubber]", "[Brand Fit=Regular, Fabric=Cotton Polyester Blend, Fabric Care=Hand wash, Fit=Regular, Ideal For=Mens, Neck Type=Round Neck, Pack of=1, Pattern=Graphic Print, Reversible=Yes]" ], 'taxo': ['', ''] } df = pd.DataFrame(data)
Step 2: Define the Extraction Function
We'll create a reusable function that handles parsing the product_details column, finding the keyword, and extracting the desired text:
def extract_target_text(df, target_keyword): # Clean up the product_details: remove brackets and split into individual key-value segments df['temp_segments'] = df['product_details'].str.strip('[]').str.split(', ') # Helper function to process each row's segments def find_and_extract(segments): for segment in segments: if target_keyword in segment: # Grab everything up to the first comma (handles edge cases where segments might have commas) return segment.split(',')[0] return None # Return None if keyword isn't found in the row # Apply the helper to each row and create a new column with results df['extracted_text'] = df['temp_segments'].apply(find_and_extract) # Optional: Remove the temporary segments column to clean up the DataFrame df = df.drop(columns='temp_segments') return df
Step 3: Use the Function
Let's test this with a keyword like slipper:
# Extract text for the keyword "slipper" result_df = extract_target_text(df, "slipper") # Print the relevant columns to verify print(result_df[['id', 'extracted_text']])
Output:
id extracted_text 0 100 Sales Package=6 Pair slipper 1 200 None
Key Notes & Adjustments
- Multiple Matches: If a row has multiple segments containing the keyword, this function returns the first match. To get all matches, modify the helper function to return a list of extracted texts instead of a single value.
- Robust Formatting: If your
product_detailshas inconsistent formatting (like extra spaces around brackets or commas), replace thestrip('[]')withstr.replace(r'^\s*\[\s*|\s*\]\s*$', '')to handle whitespace better. - Case Sensitivity: The current check is case-sensitive. To make it case-insensitive, change
if target_keyword in segment:toif target_keyword.lower() in segment.lower():.
内容的提问来源于stack exchange,提问作者outlier

