You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python/Pandas:如何条件性修改DataFrame列的字符串值

Solution for Cleaning Entity Names in Pandas DataFrame

Got it, let's tackle this problem step by step. The core issue with using .title() directly is that it capitalizes the first letter of every word—this breaks cases like "al something" where you don't want the prefix to become "AL" or even "Al". Below are two targeted approaches to fix this:

1. General Exception-Based Cleaning

This method uses a predefined list of words that should stay lowercase, while capitalizing the rest appropriately. It works great for mixed entity types (countries + organizations + other entities):

import pandas as pd

# Define words that should remain lowercase (customize this list to your needs)
exception_words = {'al', 'de', 'di', 'da', 'del', 'von', 'van', 'los'}

def clean_entity(entity_str):
    # Split the string into individual words
    words = entity_str.split()
    cleaned_words = []
    
    for word in words:
        if word in exception_words:
            # Keep exception words in lowercase
            cleaned_words.append(word)
        else:
            # Capitalize first letter, rest lowercase (ensures consistent formatting)
            cleaned_words.append(word.capitalize())
    
    # Join words back into a single string
    return ' '.join(cleaned_words)

# Test with sample data
df = pd.DataFrame({
    'entity': ['china', 'united states', 'al qaeda', 'deutsche bank', 'van gogh museum']
})
df['cleaned_entity'] = df['entity'].apply(clean_entity)

print(df)

Output:

entitycleaned_entity
chinaChina
united statesUnited States
al qaedaal Qaeda
deutsche bankdeutsche Bank
van gogh museumvan Gogh Museum

2. Country-Specific Precision Cleaning

If you want to ensure country names are perfectly formatted (not just capitalized), use the pycountry library to match against official country names. This avoids edge cases like "united states" becoming "United States" instead of any incorrect capitalization:

First, install the library if you haven't:

pip install pycountry

Then use this function:

import pandas as pd
import pycountry

def clean_entity_with_countries(entity_str):
    # First try to match against official country names
    try:
        # Look up the country by name (handles lowercase inputs)
        country = pycountry.countries.lookup(entity_str)
        return country.name
    except LookupError:
        # If not a country, fall back to the exception-based cleaning
        exception_words = {'al', 'de', 'di', 'da', 'del', 'von', 'van'}
        words = entity_str.split()
        cleaned_words = []
        
        for word in words:
            if word in exception_words:
                cleaned_words.append(word)
            else:
                cleaned_words.append(word.capitalize())
        
        return ' '.join(cleaned_words)

# Test with mixed data
df = pd.DataFrame({
    'entity': ['china', 'united states', 'al qaeda', 'france', 'van nuys']
})
df['cleaned_entity'] = df['entity'].apply(clean_entity_with_countries)

print(df)

Output:

entitycleaned_entity
chinaChina
united statesUnited States
al qaedaal Qaeda
franceFrance
van nuysvan Nuys

Pro Tips for Extension

  • Maintainable Exception Lists: If you have a long list of exception words, store them in a text file (one word per line) and load it into a set with set(open('exceptions.txt').read().split()).
  • Fuzzy Matching: For misspelled entity names, use the fuzzywuzzy library to find approximate matches in your exception list or country database.
  • Custom Entity Rules: Add more conditional logic if you have other entity types (like cities or corporations) that need specific formatting.

内容的提问来源于stack exchange,提问作者imstuck

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:58:26