Pandas新手求助:从指定列提取符合特定模式的字符串
Hey there! Since you're new to Pandas, let's walk through how to extract those specific patterns from your 'Alabama[edit]' column. From your sample data, each entry follows a pattern like Town (University Name)[citation]—so let's cover the most common extraction tasks you might need:
If you want to pull out the town name (e.g., "Auburn", "Florence") from each entry, use Pandas' str.extract() method with a regular expression. This regex captures everything from the start of the string up to the first opening parenthesis:
import pandas as pd # Assume your dataset is stored in a DataFrame called df df['Town'] = df['Alabama[edit]'].str.extract(r'^([^\(]+)').str.strip()
Breakdown of the regex and code:
r'^([^\(]+)':^Anchors the match to the start of the string.[^\(]+Matches any sequence of characters that aren't an opening parenthesis(—this stops exactly at the start of the university's parentheses.- The parentheses
()around this pattern tell Pandas to capture this part as a separate value.
.str.strip()Removes any trailing whitespace (like the space right before the parenthesis) to clean up the town name.
To get the university name inside the parentheses (e.g., "Auburn University", "University of North Alabama"), use this regex to target text between the first pair of parentheses:
df['University'] = df['Alabama[edit]'].str.extract(r'\(([^)]+)\)').str.strip()
Breakdown:
r'\(([^)]+)\)':\(Matches the opening parenthesis (we escape it with\because parentheses have special meaning in regex).([^)]+)Captures any sequence of characters that aren't a closing parenthesis).\)Matches the closing parenthesis.
.str.strip()Cleans up any extra spaces inside the parentheses.
You can do this in one step by defining two capture groups in your regex. Pandas will automatically create two new columns for the results:
# Extract both values at once df[['Town', 'University']] = df['Alabama[edit]'].str.extract(r'^([^\(]+)\(([^)]+)\)') # Clean up any extra whitespace in the new columns df['Town'] = df['Town'].str.strip() df['University'] = df['University'].str.strip()
Quick Note on Truncated Entries
For truncated entries like "Troy (Troy Unive...", the regex will still work as long as the opening parenthesis is present—it just captures whatever text is available up to the end of the string (since there's no closing parenthesis). If you have the full dataset, this won't be an issue!
内容的提问来源于stack exchange,提问作者bariumdose

