拆分DataFrame的contact列:新增性别、电话、邮箱列遇阻求助
Got it, let's work through this problem together! I'll walk you through how to clean up your contact column and extract the three new columns you need—even handling those annoying spaces after (F) or (M).
Step 1: Extract and Normalize Gender
First, we'll pull out the gender identifier and convert (M)/(F) to Male/Female. We'll use regex to target the parentheses and their contents, ignoring any surrounding spaces:
import pandas as pd # Replace this with your actual DataFrame name df = pd.DataFrame(...) # Extract gender code and map to full names df['gender'] = df['contact'].str.extract(r'\((M|F)\)')[0] df['gender'] = df['gender'].map({'M': 'Male', 'F': 'Female'})
Step 2: Clean Up Contact Content
Next, we'll remove the gender tags (and any spaces around them) from the contact column, leaving only phone numbers and/or emails:
# Remove gender markers and extra spaces, strip leading/trailing whitespace df['cleaned_contact'] = df['contact'].str.replace(r'\s*\((M|F)\)\s*', ' ', regex=True).str.strip()
Step 3: Extract Phone Number and Email
Now we can use regex to separate phone numbers and emails from the cleaned content. We'll target phone numbers (allowing for common formats like +, -, or spaces) and standard email patterns:
# Extract phone number (adjust regex if your numbers have special formats like parentheses) df['phone_number'] = df['cleaned_contact'].str.extract(r'(\+?\d[\d\-\s]*)') # Trim extra spaces from phone numbers df['phone_number'] = df['phone_number'].str.strip() # Extract email address using standard email regex df['email'] = df['cleaned_contact'].str.extract(r'([a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,})')
Step 4: (Optional) Clean Up Temporary Column
If you don't need the cleaned_contact column anymore, you can drop it:
df = df.drop('cleaned_contact', axis=1)
Example Test Case
Let's test this with sample data that includes all your edge cases:
sample_data = { 'contact': [ '(M) 1234567890', '(F) jane@example.com', # Space after (F) '9876543210 bob@test.com', '(M) 555-123-4567 alice@smith.org', '(F) 67890 charlie@doe.net' # Multiple spaces after (F) ] } df = pd.DataFrame(sample_data)
After running the code above, your DataFrame will have the three new columns correctly populated, with no issues from the spaces around gender tags.
Notes
- If your phone numbers have other formats (like
(123) 456-7890), adjust the phone regex to match—for example, user'(\+?\(?\d{3}\)?[\d\-\s]*)'for US-style numbers. - The email regex covers most standard email formats, but if you have edge cases (like country-code TLDs with longer extensions), you can tweak it as needed.
内容的提问来源于stack exchange,提问作者Yun Tae Hwang

