如何移除Pandas DataFrame某列字符串中的部分内容?
Got it, let's sort this out. Looking at your df['location.display_name'] entries, most follow a City, Region format (like Manchester, Greater Manchester) while some are just standalone city names (Cardiff, Southampton). It sounds like you want to strip out the trailing region part and keep only the city name—here's a clean, efficient way to do this with pandas:
Step-by-Step Solution
We'll use pandas' string manipulation methods to split each entry at the first comma, then grab the first part (the city) and clean up any extra whitespace:
# Create a new column with just the city name df['city'] = df['location.display_name'].str.split(',', n=1).str[0].str.strip()
Breakdown of the Code
str.split(',', n=1): Splits each string only once at the first comma. This avoids issues with entries that have special characters (like Perth, Perth & Kinross) where splitting more than once would break things..str[0]: Takes the first element from the split result (the city name before the comma). For entries without a comma (like Cardiff), this just returns the original string..str.strip(): Removes any leading/trailing whitespace that might be left after splitting (e.g., the space after the comma in Reading, Berkshire).
Example Output
Applying this to your sample data would give you:
- Kelso, Scottish Borders → Kelso
- Cardiff → Cardiff
- Hull, East Riding Of Yorkshire → Hull
- Leamington Spa, Warwickshire → Leamington Spa
If you had a different goal (like removing specific regions instead of all trailing parts), let me know and we can adjust the approach!
内容的提问来源于stack exchange,提问作者snow_fall

