You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas分割DataFrame报错:'Series'对象无'startswith'属性

修复Pandas中Series没有startswith属性的报错问题

数据背景:info.csv结构及示例

device_id,upload_time,latitude,longitude,mileage,other_vals,speed,upload_time_add_8hour,upload_time_year_month,car_id,car_type,car_num,marketer_name
1234567890123,2020-09-27 02:41:02+00:00,38.01946,114.425888,0,,0,2020/9/27 10:41,202009,,,12,
17100000001,2020-09-27 02:41:01+00:00,38.01946,114.425888,0,,0,2020/9/27 10:41,202009,,,12345,
17200000002,2020-09-25 13:46:38+00:00,38.01946,114.425888,0,,0,2020/9/25 21:46,202009,,,123456,
14111111111,2020-09-25 11:18:54+00:00,38.01946,114.425888,0,,0,2020/9/25 19:18,202009,,,12121212,
57c1e18249727a0b,2020-09-25 11:18:42+00:00,38.01946,114.425888,0,,0,2020/9/25 19:18,202009,,,,
57c1e18249727a0b,2020-09-23 10:16:55+00:00,38.01946,114.425888,0.055559317,,0,2020/9/23 18:16,202009,,,,
57c1e18249727a0b,2020-09-23 10:16:15+00:00,38.01946,114.425888,0.055559317,,0,2020/9/23 18:16,202009,,,,
57c1e18249727a0b,2020-09-23 10:15:35+00:00,38.01946,114.425888,0.055559317,,0,2020/9/23 18:15,202009,,,,
57c1e18249727a0b,2020-09-23 10:15:04+00:00,38.01946,114.425888,0.055559317,,0,2020/9/23 18:15,202009,,,,
57c1e18249727a0b,2020-09-23 10:14:55+00:00,38.01946,114.425888,0.055559317,,3.304916399,2020/9/23 18:14,202009,,,,

报错代码

import pandas as pd
df = pd.read_csv(r'info.csv', encoding='utf-8')
df_1 = df[df['device_id'].astype(str).map(len) !=11]
df_2 = df[df['device_id'].astype(str).map(len)==11 & df['device_id'].astype(str).startswith('17')]#device_id以17开头
df_3 = df[df['device_id'].astype(str).map(len)==11 & ~df['device_id'].astype(str).startswith('17')] #device_id不以17开头
df = df[pd.notnull(df['car_num'])]
print(len(df_1))
print(len(df_2))
print(len(df_3))

问题

运行上述代码后出现错误:AttributeError: 'Series' object has no attribute 'startswith',需要修复这个问题。


解决方案

这个报错有两个核心原因,我会逐一说明并给出修复方案:

1. 错误原因1:Series不能直接调用原生字符串方法

startswith是Python单个字符串的原生方法,而你这里是对Series对象调用它——Pandas的Series并没有这个属性。要对Series中的每一个字符串元素批量调用startswith,必须使用Pandas提供的str访问器,也就是str.startswith()。

2. 错误原因2:逻辑运算优先级问题

Python中&的优先级比==高,所以原代码里的len==11 & ...会被优先计算11 & ...,这完全违背了你的筛选逻辑。必须把每个独立的比较条件用括号括起来,再进行逻辑与运算。

修复后的完整代码

import pandas as pd
df = pd.read_csv(r'info.csv', encoding='utf-8')

# 修复条件的括号包裹和str访问器调用
df_1 = df[df['device_id'].astype(str).map(len) != 11]
df_2 = df[(df['device_id'].astype(str).map(len) == 11) & (df['device_id'].astype(str).str.startswith('17'))]
df_3 = df[(df['device_id'].astype(str).map(len) == 11) & (~df['device_id'].astype(str).str.startswith('17'))]

df = df[pd.notnull(df['car_num'])]
print(len(df_1))
print(len(df_2))
print(len(df_3))

额外优化建议

你可以把device_id转成字符串的操作提前执行一次,避免重复计算,让代码更高效整洁:

import pandas as pd
df = pd.read_csv(r'info.csv', encoding='utf-8')

# 提前转换device_id为字符串,减少重复操作
device_id_str = df['device_id'].astype(str)
df_1 = df[device_id_str.map(len) != 11]
df_2 = df[(device_id_str.map(len) == 11) & (device_id_str.str.startswith('17'))]
df_3 = df[(device_id_str.map(len) == 11) & (~device_id_str.str.startswith('17'))]

df = df[pd.notnull(df['car_num'])]
print(len(df_1))
print(len(df_2))
print(len(df_3))

这样修改后,代码就能正常运行,正确分割出三个子DataFrame并输出它们的长度了。

内容的提问来源于stack exchange,提问作者user9270170

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 17:07:40