You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在pandas DataFrame中忽略中间字符筛选含指定字符串的记录

实现方案

核心思路

要匹配按顺序出现A、B、C、D且四个字符之间允许插入任意字符的字符串,用正则表达式就能快速实现,对应正则规则为r'A.*?B.*?C.*?D',其中.*?为非贪婪匹配,用来匹配两个目标字符之间的任意内容。

基础实现代码

import pandas as pd

# 读取Excel文件,大文件可添加dtype参数指定列类型降低内存占用
df = pd.read_excel("你的文件路径.xlsx")

# 筛选符合要求的记录
filtered_df = df[df['col_2'].str.contains(r'A.*?B.*?C.*?D', na=False)]

# 导出筛选结果
filtered_df.to_excel("筛选结果.xlsx", index=False)

代码说明

  • na=False参数用于处理col_2为空值的场景,避免空值触发布尔索引报错
  • 如果需要忽略大小写匹配(比如同时匹配abcd、AbCd等),可以在str.contains中添加case=False参数
  • 若Excel文件体量远超内存容量,可采用分块读取的方式逐块筛选,避免内存溢出,示例如下:
import pandas as pd

chunk_result = []
# 每次读取10000行,可根据自身内存大小调整参数
for chunk in pd.read_excel("你的文件路径.xlsx", chunksize=10000):
    filtered_chunk = chunk[chunk['col_2'].str.contains(r'A.*?B.*?C.*?D', na=False)]
    chunk_result.append(filtered_chunk)

# 合并所有筛选后的块
filtered_df = pd.concat(chunk_result, ignore_index=True)

内容的提问来源于stack exchange,提问作者Chandan N

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 19:09:03