You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效在Pandas列中搜索字符串?优化慢查询代码方案

Optimizing Your Similar Keyword Search for Large Pandas DataFrames

Hey there! Let's tackle those slowdowns you're seeing as your DataFrame grows—most of the bottleneck comes from inefficient regex searches and redundant loops, so we can fix this with some targeted tweaks. Here's how to optimize your code step by step:

1. Cut Down on Redundant Combinations

Your current code loops through all possible combination lengths, then filters for 2-3 elements. Instead, we can directly generate only the combinations we need, which reduces unnecessary iterations right off the bat:

import itertools

# Generate only 2 and 3-element combinations directly
coursat = list(itertools.combinations(recommended_keywords, 2)) + list(itertools.combinations(recommended_keywords, 3))

Even better—since combinations like (a,b) and (b,a) will return the same results in your regex, we can deduplicate these to avoid repeating identical searches. Use sorted tuples to normalize pairs:

# Deduplicate 2-element pairs (since order doesn't matter for your search)
seen_pairs = set()
filtered_coursat = []
for combo in coursat:
    if len(combo) == 2:
        normalized = tuple(sorted(combo))
        if normalized not in seen_pairs:
            seen_pairs.add(normalized)
            filtered_coursat.append(combo)
    else:
        filtered_coursat.append(combo)  # Keep 3-element combos as-is
coursat = filtered_coursat

2. Replace Regex with Set Lookups (Huge Speed Boost!)

Using str.contains with dynamic regex is one of the slowest parts here—regex has to scan every string in key_words for each combo. Instead, we can pre-process your key_words column into sets (since checking membership in a set is O(1) vs. O(n) for regex):

# Preprocess key_words into a set column (do this once outside the function for repeated calls!)
df['key_words_set'] = df['key_words'].apply(set)

Then, for each combo, we can check if all elements in the combo exist in the set. This is way faster than regex:

my_list = []
for combo in coursat:
    # Filter rows where all combo keywords are present in the key_words set
    matches = df[df['key_words_set'].apply(lambda s: all(word in s for word in combo))]
    if not matches.empty:
        # Append the title values directly instead of Series to save memory/overhead
        my_list.extend(matches['title'].tolist())

If you want to avoid duplicates in the final my_list (since a single course might match multiple combos), use a set for collection:

my_set = set()
for combo in coursat:
    matches = df[df['key_words_set'].apply(lambda s: all(word in s for word in combo))]
    if not matches.empty:
        my_set.update(matches['title'].tolist())
my_list = list(my_set)

3. Optimize List Appends & Memory Usage

Instead of appending entire Series objects to my_list, extract the raw title values with .tolist() and use extend() instead of append(). This reduces the overhead of storing multiple Series and makes the final list easier to work with.

4. Bonus: Precompute Similar Keywords (If Repeated Calls)

If you're calling this function multiple times, precompute the similar keywords logic outside the function where possible, or cache results to avoid re-running the model similarity checks every time.

Full Optimized Code Example

Putting it all together, your function might look like this:

import itertools

def optimized_bla_bla(model, df):
    input_string = input("Enter your keyword string: ")
    title = input_string.split()
    titles = model.most_similar(title)
    title_list = [keyword[0] for keyword in titles]  # List comprehension instead of loop
    recommended_keywords = title_list + title

    # Generate only 2 and 3-element combinations
    coursat = list(itertools.combinations(recommended_keywords, 2)) + list(itertools.combinations(recommended_keywords, 3))

    # Deduplicate redundant 2-element pairs
    seen_pairs = set()
    filtered_coursat = []
    for combo in coursat:
        if len(combo) == 2:
            normalized = tuple(sorted(combo))
            if normalized not in seen_pairs:
                seen_pairs.add(normalized)
                filtered_coursat.append(combo)
        else:
            filtered_coursat.append(combo)
    coursat = filtered_coursat

    # Preprocess key_words to sets (do this once outside if calling multiple times!)
    if 'key_words_set' not in df.columns:
        df['key_words_set'] = df['key_words'].apply(set)

    # Collect matching titles
    matching_titles = set()
    for combo in coursat:
        matches = df[df['key_words_set'].apply(lambda s: all(word in s for word in combo))]
        if not matches.empty:
            matching_titles.update(matches['title'].tolist())

    return list(matching_titles)

These changes should drastically speed up your queries, especially as the DataFrame grows—set lookups are night-and-day faster than regex for membership checks, and cutting redundant loops reduces unnecessary work.

内容的提问来源于stack exchange,提问作者Muhammed Eltabakh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:05:20