如何清理Pandas DataFrame中以$开头的单词?
解决Pandas DataFrame中删除以$开头的单词的问题
方法一:正则表达式替换(推荐,简洁高效)
用Pandas的str.replace配合正则表达式,能快速匹配并删除所有以$开头的单词。正则r'\$\w+'会精准匹配$开头后接字母/数字的完整单词:
import pandas as pd data = ['This is awesome', '\$BTC $USD Short the market', 'Dont miss the dip on $ETH'] df = pd.DataFrame(data, columns=['text']) # 替换所有$开头的单词为空 df['cleaned_text'] = df['text'].str.replace(r'\$\w+', '', regex=True) # 清理替换后可能出现的多余空格 df['cleaned_text'] = df['cleaned_text'].str.strip().str.replace(r'\s+', ' ', regex=True) print(df['cleaned_text'])
输出结果:
0 This is awesome 1 Short the market 2 Dont miss the dip on Name: cleaned_text, dtype: object
方法二:拆分过滤法(符合你之前的思路)
如果你想通过拆分单词+startswith()实现,可以用str.split把字符串拆成单词列表,过滤掉以$开头的单词后再拼接回去:
df['cleaned_text'] = df['text'].apply(lambda x: ' '.join([word for word in x.split() if not word.startswith('$')]))
这个方法逻辑更直观,适合新手理解,但处理大数据集时效率略低于正则方法。
注意点
- 原数据中的
\$BTC是转义后的$,实际存储中如果是直接的$BTC,上述代码依然有效,因为正则里的\$会匹配字面量$。 - 处理后记得清理多余空格,避免出现多个连续空格的情况。
内容的提问来源于stack exchange,提问作者user19783276
相关产品推荐
相关产品推荐

