You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过字段子串过滤Pandas DataFrame?排查索引错误及通用方案

Pandas字符串日期字段筛选:错误排查与通用解法

一、首次代码报错原因

你写的shops['opening'][-2:] != "22"逻辑错误:

  • 直接对Series用[-2:]是对整个Series的行索引切片,取的是数据中最后2行的opening值,得到长度为2的布尔Series。
  • 原DataFrame行数远大于2,导致布尔索引长度与原数据不匹配,触发IndexingError。
  • 正确操作是对Series中每个字符串单独切片,必须用Pandas的.str访问器,即shops['opening'].str[-2:]。

二、通用简便筛选方法

1. 直接用.str切片(最直观)

和Python字符串切片规则一致,可快速定位任意位置子串:

# 筛选最后两位不是"22"的数据
filtered_shops = shops[shops['opening'].str[-2:] != "22"]

# 示例:筛选前四位是"2024"的数据
filtered_shops = shops[shops['opening'].str[:4] == "2024"]

# 示例:筛选第3到第5位(索引2到4,左闭右开)是"061"的数据
filtered_shops = shops[shops['opening'].str[2:5] == "061"]

2. 用.str.slice()方法(更规范)

功能和.str[]一致,适合明确指定起始/结束位置的场景:

# 取最后两位,等价于.str[-2:]
filtered_shops = shops[shops['opening'].str.slice(start=-2) != "22"]

# 取第2到第4位(索引1到3)
filtered_shops = shops[shops['opening'].str.slice(start=1, stop=4) == "123"]

3. 正则匹配(最灵活,适用于复杂模式)

如果需要复杂匹配规则(比如任意位置包含特定字符、自定义位置匹配),用.str.contains()配合正则:

# 筛选结尾不是"22"的数据(等价于endswith('22')取反)
filtered_shops = shops[~shops['opening'].str.contains(r'22$')]

# 筛选任意位置包含"10"的数据
filtered_shops = shops[shops['opening'].str.contains(r'10')]

# 筛选第5位是"5"的数据(^表示开头,.匹配任意单个字符)
filtered_shops = shops[shops['opening'].str.contains(r'^....5')]

内容的提问来源于stack exchange,提问作者Félix Rodriguez Moya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 00:32:17