如何移除DataFrame的CityIds列中长度小于4的项或单独拆分?
处理DataFrame中CityIds列的短项过滤与拆分
针对你的需求,这里提供两种高效的处理方案(适配大数据量场景):
方案一:删除长度小于4的项,保留符合长度要求的内容
通过正则拆分字符串(自动处理逗号后的空格),过滤出长度≥4的元素后重新合并:
import pandas as pd # 构造示例DataFrame(实际使用时替换为你的数据) df = pd.DataFrame({ 'CityIds': [ '98765, 98-oki, th6, iuy89, 8.90765', '89ol, gh98.0p, klopi, th, loip', '98087,PAKJIYT, hju, yu8oi, iupli' ] }) # 生成过滤后的列 df['Filtered_CityIds'] = df['CityIds'].str.split(r',\s*').apply( lambda x: ', '.join([item for item in x if len(item) >= 4]) )
处理后结果:
| CityIds | Filtered_CityIds |
|---|---|
| 98765, 98-oki, th6, iuy89, 8.90765 | 98765, 98-oki, iuy89, 8.90765 |
| 89ol, gh98.0p, klopi, th, loip | 89ol, gh98.0p, klopi, loip |
| 98087,PAKJIYT, hju, yu8oi, iupli | 98087, PAKJIYT, yu8oi, iupli |
方案二:将短项拆分到单独列
把长度<4的项和符合要求的项分别放到两个列中:
# 定义拆分函数 def split_short_long(items): short_items = [item for item in items if len(item) < 4] long_items = [item for item in items if len(item) >= 4] return ', '.join(long_items), ', '.join(short_items) # 应用函数生成两列 df[['Long_CityIds', 'Short_CityIds']] = df['CityIds'].str.split(r',\s*').apply( lambda x: pd.Series(split_short_long(x)) )
处理后结果:
| CityIds | Long_CityIds | Short_CityIds |
|---|---|---|
| 98765, 98-oki, th6, iuy89, 8.90765 | 98765, 98-oki, iuy89, 8.90765 | th6 |
| 89ol, gh98.0p, klopi, th, loip | 89ol, gh98.0p, klopi, loip | th |
| 98087,PAKJIYT, hju, yu8oi, iupli | 98087, PAKJIYT, yu8oi, iupli | hju |
说明
- 用
str.split(r',\s*')通过正则匹配逗号(后面可跟任意空格),确保拆分不受空格影响; - 采用
apply结合列表推导式处理,比纯Python循环效率更高,适合大数据量场景; - 如果你的DataFrame数据量极大,可以考虑用
swifter库加速apply操作(需额外安装)。
内容的提问来源于stack exchange,提问作者Luke Lucy
相关产品推荐
相关产品推荐

