Python/Pandas:如何条件性修改DataFrame列的字符串值
Got it, let's tackle this problem step by step. The core issue with using .title() directly is that it capitalizes the first letter of every word—this breaks cases like "al something" where you don't want the prefix to become "AL" or even "Al". Below are two targeted approaches to fix this:
1. General Exception-Based Cleaning
This method uses a predefined list of words that should stay lowercase, while capitalizing the rest appropriately. It works great for mixed entity types (countries + organizations + other entities):
import pandas as pd # Define words that should remain lowercase (customize this list to your needs) exception_words = {'al', 'de', 'di', 'da', 'del', 'von', 'van', 'los'} def clean_entity(entity_str): # Split the string into individual words words = entity_str.split() cleaned_words = [] for word in words: if word in exception_words: # Keep exception words in lowercase cleaned_words.append(word) else: # Capitalize first letter, rest lowercase (ensures consistent formatting) cleaned_words.append(word.capitalize()) # Join words back into a single string return ' '.join(cleaned_words) # Test with sample data df = pd.DataFrame({ 'entity': ['china', 'united states', 'al qaeda', 'deutsche bank', 'van gogh museum'] }) df['cleaned_entity'] = df['entity'].apply(clean_entity) print(df)
Output:
| entity | cleaned_entity |
|---|---|
| china | China |
| united states | United States |
| al qaeda | al Qaeda |
| deutsche bank | deutsche Bank |
| van gogh museum | van Gogh Museum |
2. Country-Specific Precision Cleaning
If you want to ensure country names are perfectly formatted (not just capitalized), use the pycountry library to match against official country names. This avoids edge cases like "united states" becoming "United States" instead of any incorrect capitalization:
First, install the library if you haven't:
pip install pycountry
Then use this function:
import pandas as pd import pycountry def clean_entity_with_countries(entity_str): # First try to match against official country names try: # Look up the country by name (handles lowercase inputs) country = pycountry.countries.lookup(entity_str) return country.name except LookupError: # If not a country, fall back to the exception-based cleaning exception_words = {'al', 'de', 'di', 'da', 'del', 'von', 'van'} words = entity_str.split() cleaned_words = [] for word in words: if word in exception_words: cleaned_words.append(word) else: cleaned_words.append(word.capitalize()) return ' '.join(cleaned_words) # Test with mixed data df = pd.DataFrame({ 'entity': ['china', 'united states', 'al qaeda', 'france', 'van nuys'] }) df['cleaned_entity'] = df['entity'].apply(clean_entity_with_countries) print(df)
Output:
| entity | cleaned_entity |
|---|---|
| china | China |
| united states | United States |
| al qaeda | al Qaeda |
| france | France |
| van nuys | van Nuys |
Pro Tips for Extension
- Maintainable Exception Lists: If you have a long list of exception words, store them in a text file (one word per line) and load it into a set with
set(open('exceptions.txt').read().split()). - Fuzzy Matching: For misspelled entity names, use the
fuzzywuzzylibrary to find approximate matches in your exception list or country database. - Custom Entity Rules: Add more conditional logic if you have other entity types (like cities or corporations) that need specific formatting.
内容的提问来源于stack exchange,提问作者imstuck

