You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python优化:超大规模列表中匹配字符串条件的索引查找提速方法

优化站点组合索引查询性能方案

可复现代码

## Mock Data ##
my_list = list(range(1700))

import itertools
cross_product = list(itertools.product(my_list, my_list))
station_combinations = ["_".join([str(i), str(b)]) for i, b in cross_product if i != b]
###############

from time import time
station_name = "5"

start = time()

for h in range(10):
    reverse_indexes = [count for count, j in enumerate(station_combinations) if j.split("_")[1] == station_name]
    regular_indexes = [count for count, j in enumerate(station_combinations) if j.split("_")[0] == station_name]

print(f"原代码耗时: {time() - start:.2f} 秒")

背景说明

station_combinations是my_list的笛卡尔积(排除i=b的情况),元素格式为a_b字符串,代表从站点a到b的路径:

  • reverse_indexes:所有目的地为指定站点(b等于station_name)的元素索引
  • regular_indexes:所有出发地为指定站点(a等于station_name)的元素索引

问题描述

现有代码运行速度极慢,10次循环耗时约8秒,实际场景需循环约2000次。尝试过numba但因DataFrame数据无法适配@njit,需寻求显著提速方案。

优化方案

核心思路是预先构建索引映射字典,避免每次循环重复遍历、拆分字符串的冗余操作:

步骤1:预先构建全局映射

在循环执行前,一次性遍历station_combinations,构建两个映射字典:

  • regular_map:键为出发站字符串,值为对应所有元素的索引列表
  • reverse_map:键为到达站字符串,值为对应所有元素的索引列表

代码实现:

from time import time

# 预先构建映射(仅执行一次)
regular_map = {}
reverse_map = {}

for idx, combo in enumerate(station_combinations):
    a_str, b_str = combo.split("_")
    # 更新出发站映射
    if a_str not in regular_map:
        regular_map[a_str] = []
    regular_map[a_str].append(idx)
    # 更新到达站映射
    if b_str not in reverse_map:
        reverse_map[b_str] = []
    reverse_map[b_str].append(idx)

# 测试查询效率
station_name = "5"
start = time()

for h in range(10):
    reverse_indexes = reverse_map.get(station_name, [])
    regular_indexes = regular_map.get(station_name, [])

print(f"优化后代码耗时: {time() - start:.6f} 秒")

优化原理

  • 原代码每次循环需遍历约2.89M个元素并拆分字符串,10次循环累计执行近29M次冗余操作
  • 优化后仅需一次遍历构建映射,后续每次查询直接通过字典取值,时间复杂度从O(n)降至O(1),2000次循环的耗时可忽略不计

适配DataFrame场景

若station_combinations来自DataFrame的某一列,可利用pandas分组功能构建映射,避免手动遍历:

import pandas as pd

# 假设df是包含station_combinations列的DataFrame
df = pd.DataFrame({"combo": station_combinations})
df[["a", "b"]] = df["combo"].str.split("_", expand=True)

# 构建出发站到索引的映射
regular_map_df = df.groupby("a")["combo"].apply(lambda x: x.index.tolist()).to_dict()
# 构建到达站到索引的映射
reverse_map_df = df.groupby("b")["combo"].apply(lambda x: x.index.tolist()).to_dict()

# 查询示例
reverse_indexes = reverse_map_df.get(station_name, [])
regular_indexes = regular_map_df.get(station_name, [])

内容的提问来源于stack exchange,提问作者Cindy Burker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 02:35:24