You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何清理Pandas DataFrame中以$开头的单词?

解决Pandas DataFrame中删除以$开头的单词的问题

方法一:正则表达式替换(推荐,简洁高效)

用Pandas的str.replace配合正则表达式,能快速匹配并删除所有以$开头的单词。正则r'\$\w+'会精准匹配$开头后接字母/数字的完整单词:

import pandas as pd
data = ['This is awesome', '\$BTC $USD Short the market', 'Dont miss the dip on $ETH']
df = pd.DataFrame(data, columns=['text'])

# 替换所有$开头的单词为空
df['cleaned_text'] = df['text'].str.replace(r'\$\w+', '', regex=True)
# 清理替换后可能出现的多余空格
df['cleaned_text'] = df['cleaned_text'].str.strip().str.replace(r'\s+', ' ', regex=True)

print(df['cleaned_text'])

输出结果:

0                This is awesome
1             Short the market
2    Dont miss the dip on
Name: cleaned_text, dtype: object

方法二:拆分过滤法(符合你之前的思路)

如果你想通过拆分单词+startswith()实现,可以用str.split把字符串拆成单词列表,过滤掉以$开头的单词后再拼接回去:

df['cleaned_text'] = df['text'].apply(lambda x: ' '.join([word for word in x.split() if not word.startswith('$')]))

这个方法逻辑更直观,适合新手理解,但处理大数据集时效率略低于正则方法。

注意点

  • 原数据中的\$BTC是转义后的$,实际存储中如果是直接的$BTC,上述代码依然有效,因为正则里的\$会匹配字面量$。
  • 处理后记得清理多余空格,避免出现多个连续空格的情况。

内容的提问来源于stack exchange,提问作者user19783276

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 22:24:22