You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas Series中实现无序工时内容的匹配搜索?

Pandas工时组合匹配:忽略顺序实现匹配

问题场景

现有代码仅能在输入工时字符串与Pandas Series中的内容完全顺序匹配时返回对应日期,当输入的工时顺序打乱(如输入04 04 22 12 06 09),无法匹配到实际对应的03/10/2023日期。需要实现不考虑工时顺序的组合匹配功能。

原代码如下:

num_srch = "12 22 04 06 09 04"
num_mask = pwrDF['Work Hours'] == num_srch
num_match = pwrDF.loc[num_mask, "Work Date"].values

if len(pwrDF[(pwrDF['Work Hours'] == num_srch)]) >0:
    num_match = num_match[0]
    print('These hours match this date:\n ', num_match)
else:
    print('No hours match what you entered')

对应的DataFrame数据:

Work DateWork Hours
003/10/202312 22 04 06 09 04
103/09/202302 13 29 58 69 05
203/08/202303 14 05 08 09 06

解决方案

核心思路是对工时字符串做标准化处理:将字符串拆分、排序后重新拼接,让相同元素组合(不管顺序)得到一致的标准化结果,再通过标准化结果进行匹配。

实现代码

import pandas as pd

# 初始化示例DataFrame(可替换为你的实际数据)
data = {
    "Work Date": ["03/10/2023", "03/09/2023", "03/08/2023"],
    "Work Hours": ["12 22 04 06 09 04", "02 13 29 58 69 05", "03 14 05 08 09 06"]
}
pwrDF = pd.DataFrame(data)

# 用户输入的工时(顺序打乱)
num_srch = "04 04 22 12 06 09"

# 定义标准化函数:拆分字符串→排序→重新拼接
def standardize_hours(hours_str):
    sorted_parts = sorted(hours_str.split())
    return ' '.join(sorted_parts)

# 标准化搜索字符串
standardized_search = standardize_hours(num_srch)

# 匹配标准化后的工时,获取对应日期
matches = pwrDF[pwrDF['Work Hours'].apply(standardize_hours) == standardized_search]['Work Date']

if not matches.empty:
    print(f'These hours match this date:\n  {matches.iloc[0]}')
else:
    print('No hours match what you entered')

代码说明

  1. 标准化函数:standardize_hours将任意顺序的工时字符串转换为排序后的标准格式,例如"12 22 04 06 09 04"和"04 04 22 12 06 09"标准化后都会变成"04 04 06 09 12 22"。
  2. 匹配逻辑:通过apply对DataFrame的Work Hours列批量标准化,再与搜索字符串的标准化结果对比,直接筛选出匹配的日期。
  3. 优化点:无需新增临时列,直接在匹配时处理,代码更简洁高效。

内容的提问来源于stack exchange,提问作者VJ1222

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 05:42:48