You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取文本时间参照并按提取时间升序排序对应数据行

字符串时间提取与排序实现方案

核心逻辑

将不同单位的相对时间统一转换为可比较的数值维度(比如秒级时间差、实际发生时间戳),再基于该维度排序即可解决通用时间适配的问题,不需要针对不同单位单独写判断逻辑,新增时间单位只要更新映射表即可。

实现步骤

  • 预定义时间单位映射:把所有支持的时间单位换算为统一的最小单位(比如秒),后续扩展单位直接更新映射表即可,适配性强。
# 时间单位转秒的映射,可按需扩展分钟、月、年等单位
UNIT_MAP = {
    'hour': 3600,
    'day': 86400,
    'week': 604800
}
  • 提取时间片段:按你提到的分隔符「—」分割Text字段,取前半部分作为待匹配的时间片段,无分隔符或者Text为空直接判定为无有效时间。
  • 正则匹配时间信息:用正则表达式从时间片段中提取数值和单位,匹配失败则归为无有效时间组。
import re
# 自动适配单复数,比如hour/hours都可以匹配
time_pattern = re.compile(r'(\d+)\s*(' + '|'.join(UNIT_MAP.keys()) + r')s?\s*ago')
  • 统一转换为可比较值:把匹配到的数值乘上对应单位的换算值,得到总秒数,数值越大代表距离当前时间越久;无有效时间的赋值为无穷大,排序时自动排在最后。
  • 排序并生成排序编号:按转换后的秒数排序,相同秒数的记录使用相同的排序编号,无有效时间的记录不填排序编号。如果需要和你给出的示例排序规则一致,调整排序方向或者编号生成逻辑即可。

完整可运行示例代码

import re
from math import inf

# 配置项
UNIT_MAP = {
    'hour': 3600,
    'day': 86400,
    'week': 604800
}
time_pattern = re.compile(r'(\d+)\s*(' + '|'.join(UNIT_MAP.keys()) + r')s?\s*ago')

# 示例数据
data = [
    {"Customer": 1, "Text": "12 hours ago — the customer applied for a discount"},
    {"Customer": 2, "Text": "6 hours ago — the customer contacted the customer service"},
    {"Customer": 3, "Text": "1 day ago — the customer reported an issue"},
    {"Customer": 4, "Text": "1 day ago — no answer"},
    {"Customer": 4, "Text": "2 days ago — Open issue"},
    {"Customer": 5, "Text": ""}
]

# 给每条数据计算排序权重
for item in data:
    text = item["Text"].strip()
    weight = inf
    if "—" in text:
        time_part = text.split("—")[0].strip()
        match_res = time_pattern.search(time_part)
        if match_res:
            num = int(match_res.group(1))
            unit = match_res.group(2)
            weight = num * UNIT_MAP[unit]
    item["sort_weight"] = weight

# 按权重升序排序(时间越近排在越前面,要调整顺序把reverse改成True即可)
sorted_data = sorted(data, key=lambda x: x["sort_weight"])

# 生成排序编号,相同权重编号一致
current_rank = 0
prev_weight = None
for item in sorted_data:
    if item["sort_weight"] == inf:
        item["Sort by"] = ""
        continue
    if item["sort_weight"] != prev_weight:
        current_rank += 1
        prev_weight = item["sort_weight"]
    item["Sort by"] = current_rank

# 格式化输出结果
print(f"{'Customer':<10}{'Text':<70}{'Sort by'}")
for item in sorted_data:
    print(f"{item['Customer']:<10}{item['Text']:<70}{item['Sort by']}")

内容的提问来源于stack exchange,提问作者LdM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 07:06:00