如何用apply从DataFrame多列提取括号内内容
Got it, let's work through this. You mentioned extracting bracket content works perfectly with vectors, but hitting a wall when using apply on DataFrame columns—here's how to get this working reliably.
Step 1: Define a Robust Extraction Function
First, let's create a function that handles all edge cases (like missing values, non-string entries, or text without brackets) so it doesn't break when applied to your DataFrame. We'll use regex for the extraction:
import re import pandas as pd def pull_bracket_text(input_text): # Handle NaNs, numbers, or non-string values first if pd.isna(input_text) or not isinstance(input_text, str): return pd.NA # Match content inside parentheses (adjust regex for other brackets!) # Use r'\[(.*?)\]' for square brackets, r'\{(.*?)\}' for curly braces matches = re.findall(r'\((.*?)\)', input_text) # Return first match if exists, else NA return matches[0] if matches else pd.NA
Step 2: Apply the Function to Multiple Columns
You have two straightforward options depending on whether you want to process all columns or just a subset:
Option 1: Process Every Column in the DataFrame
Use applymap—it's designed for element-wise operations across the entire DataFrame, which is exactly what we need here:
# Sample DataFrame to test with df = pd.DataFrame({ 'fruit': ['Apple (Red)', 'Banana (Yellow)', 'Grape'], 'animal': ['Cat (Mammal)', pd.NA, 'Dog (Canine)'], 'vehicle': ['Car (Sedan)', '', 'Bike (Mountain)'] }) # Apply extraction to all columns df_extracted = df.applymap(pull_bracket_text)
Option 2: Process Only Specific Columns
If you don't want to touch every column, target the ones you need and use applymap on that subset:
# List columns you want to process target_columns = ['fruit', 'animal'] # Overwrite the original columns with extracted content df[target_columns] = df[target_columns].applymap(pull_bracket_text)
Common Pitfalls to Fix Why Your Original Attempt Failed
- No handling for non-string values: If your DataFrame has NaNs, integers, or floats, your original vector-based code might not account for these—our function checks for this explicitly.
- Wrong
applyaxis: If you useddf.apply(..., axis=1), that processes rows instead of individual elements. Stick toapplymapfor element-wise work, or usedf[col].apply(pull_bracket_text)for single columns. - Mismatched regex: Double-check your regex matches the bracket type you're targeting (e.g., square vs. round brackets). Our example uses round—adjust the pattern if needed.
内容的提问来源于stack exchange,提问作者Al14

