Python优化:超大规模列表中匹配字符串条件的索引查找提速方法
优化站点组合索引查询性能方案
可复现代码
## Mock Data ## my_list = list(range(1700)) import itertools cross_product = list(itertools.product(my_list, my_list)) station_combinations = ["_".join([str(i), str(b)]) for i, b in cross_product if i != b] ############### from time import time station_name = "5" start = time() for h in range(10): reverse_indexes = [count for count, j in enumerate(station_combinations) if j.split("_")[1] == station_name] regular_indexes = [count for count, j in enumerate(station_combinations) if j.split("_")[0] == station_name] print(f"原代码耗时: {time() - start:.2f} 秒")
背景说明
station_combinations是my_list的笛卡尔积(排除i=b的情况),元素格式为a_b字符串,代表从站点a到b的路径:
reverse_indexes:所有目的地为指定站点(b等于station_name)的元素索引regular_indexes:所有出发地为指定站点(a等于station_name)的元素索引
问题描述
现有代码运行速度极慢,10次循环耗时约8秒,实际场景需循环约2000次。尝试过numba但因DataFrame数据无法适配@njit,需寻求显著提速方案。
优化方案
核心思路是预先构建索引映射字典,避免每次循环重复遍历、拆分字符串的冗余操作:
步骤1:预先构建全局映射
在循环执行前,一次性遍历station_combinations,构建两个映射字典:
regular_map:键为出发站字符串,值为对应所有元素的索引列表reverse_map:键为到达站字符串,值为对应所有元素的索引列表
代码实现:
from time import time # 预先构建映射(仅执行一次) regular_map = {} reverse_map = {} for idx, combo in enumerate(station_combinations): a_str, b_str = combo.split("_") # 更新出发站映射 if a_str not in regular_map: regular_map[a_str] = [] regular_map[a_str].append(idx) # 更新到达站映射 if b_str not in reverse_map: reverse_map[b_str] = [] reverse_map[b_str].append(idx) # 测试查询效率 station_name = "5" start = time() for h in range(10): reverse_indexes = reverse_map.get(station_name, []) regular_indexes = regular_map.get(station_name, []) print(f"优化后代码耗时: {time() - start:.6f} 秒")
优化原理
- 原代码每次循环需遍历约2.89M个元素并拆分字符串,10次循环累计执行近29M次冗余操作
- 优化后仅需一次遍历构建映射,后续每次查询直接通过字典取值,时间复杂度从O(n)降至O(1),2000次循环的耗时可忽略不计
适配DataFrame场景
若station_combinations来自DataFrame的某一列,可利用pandas分组功能构建映射,避免手动遍历:
import pandas as pd # 假设df是包含station_combinations列的DataFrame df = pd.DataFrame({"combo": station_combinations}) df[["a", "b"]] = df["combo"].str.split("_", expand=True) # 构建出发站到索引的映射 regular_map_df = df.groupby("a")["combo"].apply(lambda x: x.index.tolist()).to_dict() # 构建到达站到索引的映射 reverse_map_df = df.groupby("b")["combo"].apply(lambda x: x.index.tolist()).to_dict() # 查询示例 reverse_indexes = reverse_map_df.get(station_name, []) regular_indexes = regular_map_df.get(station_name, [])
内容的提问来源于stack exchange,提问作者Cindy Burker
相关产品推荐
相关产品推荐

