You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python按条件过滤列表:现有自定义函数的优化改进咨询

Optimizing Your Condition-Based List Filter Function

Great question! Let's start by unpacking the behavior of your original function and then look at cleaner, more efficient ways to achieve the same goal (while fixing a few hidden edge cases along the way).

Core Goal Recap

Your function aims to find all items in cond_list that appear as a substring in at least one element of input_list, then return those matching conditions without duplicates.

Issues with the Original Implementation

Your current code works for the test case, but has a few areas to improve:

  • Unnecessary deduplication: Converting input_list to a set and back wastes memory and processing time—we only need to know if a condition exists in any element, not how many times it appears.
  • Duplicate result risk: If multiple input elements match the same condition, your function would append that condition to res multiple times. Your test case avoids this only because deduplication removed repeated inputs, which isn't reliable for all scenarios.
  • Verbose loops: The nested for-loop works, but we can make the code far more readable with Python's built-in tools.

Optimized Solutions

Option 1: Preserve Condition Order (Python 3.7+)

This version uses a list comprehension with any() to check for matches, then uses dict.fromkeys() to remove duplicates while keeping the original order of cond_list (dictionaries preserve insertion order in Python 3.7+):

def myFunction(cond_list, input_list):
    # Collect all conditions with at least one match in input_list
    matching_conds = [cond for cond in cond_list if any(cond in item for item in input_list)]
    # Remove duplicates while retaining original order
    return list(dict.fromkeys(matching_conds))

Option 2: Shorter, Faster (Order May Vary)

If you don't need to preserve the original order of cond_list, a set comprehension is slightly more efficient:

def myFunction(cond_list, input_list):
    return list({cond for cond in cond_list if any(cond in item for item in input_list)})

Option 3: Optimized for Large Input Lists

If your input_list is extremely large, pre-process it into a single string (with a unique separator) to reduce repeated substring checks:

def myFunction(cond_list, input_list):
    # Join inputs with a separator that won't appear in conditions/inputs
    input_str = "###SEP###".join(input_list)
    return list(dict.fromkeys([cond for cond in cond_list if cond in input_str]))

Note: Use a separator you're certain won't appear in your data to avoid false negatives.

Test with Your Sample Data

Using your test case:

cond = ['cat', 'rabbit']
input_list = ['', 'cat 88.96%', '.', 'I have a dog', '', 'rabbit 12.44%', '', 'I like tiger']
print(myFunction(cond, input_list))  # Output: ['cat', 'rabbit']

All optimized versions return the correct result, with cleaner logic and fewer edge cases.

内容的提问来源于stack exchange,提问作者vincentlai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:36:36