You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升Pandas str.contains速度?百万行DataFrame字符串模式搜索优化

Hey there! Let's tackle this problem step by step—first fixing the OR logic in your apply approach, then speeding up the operation for your million-row DataFrame.

First: Fix the OR Logic in Your Apply Method

Your current apply approach fails to recognize the | operator because in treats the entire string "home|property|house|apartment" as a single match target. To fix this, split the pattern into a set of keywords and check if any of them exist in the string:

pattern_set = set(pattern_example.lower().split("|"))
df[field].apply(lambda x: any(word in x.lower() for word in pattern_set))

Note: While this fixes the logic issue, it’s still a row-wise loop under the hood—so it won’t be as fast as vectorized operations. Let’s move on to speed optimizations.

Speed Optimizations for String Matching

For million-row datasets, vectorized operations are your best bet. Here are the most effective tweaks:

1. Precompile Your Regular Expression

str.contains recompiles the regex every time it runs, which wastes time on large datasets. Precompile once to avoid this overhead:

import re
# Precompile with case-insensitive flag to skip redundant case=False parameter
compiled_pattern = re.compile(pattern_example.lower(), re.IGNORECASE)
result = df[field].str.contains(compiled_pattern, regex=True)

This alone can cut runtime by 30-50% for large datasets.

2. Switch to Pandas' StringDtype

If your column uses the default object dtype, converting to string (Pandas' optimized string type) can boost performance:

df[field] = df[field].astype("string")

StringDtype reduces memory overhead and optimizes string operations compared to the generic object type.

3. Avoid Repeated Lowercasing

Instead of converting strings to lowercase inside str.contains every time, do it once upfront:

# Normalize the column to lowercase once
df[field] = df[field].str.lower()
# Match against the pre-lowercased pattern
result = df[field].str.contains(pattern_example.lower(), regex=True)

This eliminates redundant string transformation work during the matching step.

4. Try Vectorized numpy Operations (Alternative)

If you want to bypass Pandas' str methods, use numpy’s vectorized function with the precompiled regex:

import numpy as np
vectorized_search = np.vectorize(lambda x: compiled_pattern.search(x) is not None)
result = vectorized_search(df[field].values)

Performance is comparable to precompiled str.contains, but test both to see which works better for your data.

5. Third-Party Tools for Extreme Scale

If your dataset grows beyond a few million rows, consider these tools:

  • Swifter: Automatically chooses between vectorized operations and optimized apply for mixed data:
    import swifter
    result = df[field].swifter.apply(lambda x: compiled_pattern.search(x) is not None)
    
  • Dask: Splits your data into chunks and processes them in parallel, ideal for out-of-core datasets.

Quick Performance Rankings

From fastest to slowest for most use cases:

  1. Precompiled regex + StringDtype
  2. Precompiled regex with object dtype
  3. numpy vectorized search
  4. Fixed apply method (row-wise loop)

内容的提问来源于stack exchange,提问作者Daniel Hangan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:38:38