如何通过正则表达式移除Pandas文本列中的字母数字混合单词?
Got it, let's fix this for you! Your current regex is targeting the wrong thing—it's removing non-alphanumeric/non-space characters, which isn't what you need. What you actually want is to delete entire words that contain both letters and numbers (like c6c587469, d4a29a, etc.).
Working Pandas Code
Here's the correct regex and code to achieve your desired output:
# Replace alphanumeric hybrid words with empty string df_colum = df_colum.str.replace(r'\b(?=.*[A-Za-z])(?=.*\d)\w+\b', '', regex=True) # Optional: Clean up extra spaces left behind (multiple spaces → single space, trim edges) df_colum = df_colum.str.replace(r'\s+', ' ', regex=True).str.strip()
Regex Breakdown
Let's break down why this works:
\b: Matches a word boundary, ensuring we target full words (not parts of words)(?=.*[A-Za-z]): Positive lookahead to confirm the word contains at least one letter(?=.*\d): Positive lookahead to confirm the word contains at least one digit\w+: Matches one or more word characters (letters, digits, underscores—adjust to[A-Za-z0-9]+if you don't want to include underscores)\b: Closing word boundary to complete the full word match
Example Test
If we apply this to your input text:
due to previous assess c6c587469 and 4ec0f198 nearest and with fill station in the citi becaus of our satisfact in the d4a29a already averaging my thoughts on e977f33588f react to
After running the code, you'll get exactly your desired output:
due to previous assess and nearest and with fill station in the citi becaus of our satisfact in the already averaging my thoughts on react to
内容的提问来源于stack exchange,提问作者misterbenj34

