如何删除以特定通配符字符串结尾的DataFrame列?
Solution to Drop Specific Publication Date Columns
First, let's fix the issues in your original code:
- You tried to call
.str.contains()on individual string column names (c), which isn't valid—.stris a pandas accessor for Series/Index objects, not single strings. - Your regex pattern
[a+z]was incorrect; it should be[a-z]to match any lowercase letter (the+was a typo here).
Here are two working approaches to achieve your goal:
Approach 1: Using Pandas' Built-in String Methods
This is the most concise way, leveraging pandas' Index.str methods to filter columns:
# Filter columns that match the pattern: publication_date_ followed by a lowercase letter at the end columns_to_drop = df_merge.columns[df_merge.columns.str.contains(r'publication_date_[a-z]$')] # Drop the identified columns df_merge = df_merge.drop(columns_to_drop, axis=1)
Regex Breakdown:
publication_date_: Exact match for the prefix[a-z]: Matches any single lowercase letter$: Ensures the string ends right after the letter (so no extra characters are present)
Approach 2: Using the re Module for Explicit Matching
If you prefer using Python's standard regex library:
import re # Iterate through columns and collect those matching the pattern columns_to_drop = [col for col in df_merge.columns if re.fullmatch(r'publication_date_[a-z]', col)] # Drop the columns df_merge = df_merge.drop(columns_to_drop, axis=1)
Why This Works:
re.fullmatch()checks if the entire column name matches the pattern, which is equivalent to adding$in the previous regex.- This will target exactly columns like
publication_date_xandpublication_date_y, leavingpublication_date(without the suffix) intact.
After running either of these, your DataFrame will have the desired columns: contract_number, award_date, award_date_x, award_date_y, publication_date.
内容的提问来源于stack exchange,提问作者Mitchell
相关产品推荐
相关产品推荐

