You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过正则表达式移除Pandas文本列中的字母数字混合单词?

Solution to Remove Alphanumeric Hybrid Words in Pandas

Got it, let's fix this for you! Your current regex is targeting the wrong thing—it's removing non-alphanumeric/non-space characters, which isn't what you need. What you actually want is to delete entire words that contain both letters and numbers (like c6c587469, d4a29a, etc.).

Working Pandas Code

Here's the correct regex and code to achieve your desired output:

# Replace alphanumeric hybrid words with empty string
df_colum = df_colum.str.replace(r'\b(?=.*[A-Za-z])(?=.*\d)\w+\b', '', regex=True)

# Optional: Clean up extra spaces left behind (multiple spaces → single space, trim edges)
df_colum = df_colum.str.replace(r'\s+', ' ', regex=True).str.strip()

Regex Breakdown

Let's break down why this works:

  • \b: Matches a word boundary, ensuring we target full words (not parts of words)
  • (?=.*[A-Za-z]): Positive lookahead to confirm the word contains at least one letter
  • (?=.*\d): Positive lookahead to confirm the word contains at least one digit
  • \w+: Matches one or more word characters (letters, digits, underscores—adjust to [A-Za-z0-9]+ if you don't want to include underscores)
  • \b: Closing word boundary to complete the full word match

Example Test

If we apply this to your input text:

due to previous assess c6c587469 and 4ec0f198 nearest and with fill station in the citi becaus of our satisfact in the d4a29a already averaging my thoughts on e977f33588f react to

After running the code, you'll get exactly your desired output:

due to previous assess and nearest and with fill station in the citi becaus of our satisfact in the already averaging my thoughts on react to

内容的提问来源于stack exchange,提问作者misterbenj34

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:31:52