You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

表格汉字匹配分组问题求助:已完成目标1,求目标2实现方案

匹配列表(Matchlist)分组任务完整实现

需求回顾

  • 第一组:与匹配列表精确匹配的目标词
  • 第二组:未在匹配列表中的目标词,且匹配列表存在以该目标词首字符开头的条目,需将目标词与对应匹配条目关联
  • 第三组:未在匹配列表中,且匹配列表无对应首字符开头条目的目标词

修正后的目标1实现代码

你原来的代码存在逻辑颠倒(应检查目标词是否在匹配列表中,而非反过来)和语法问题(中文引号、Yes/No未加引号),修正后如下:

import pandas as pd

# 假设table是包含'Match list'和'Target word'列的DataFrame
match_series = table['Match list'].dropna().unique()  # 去重并清理空值,作为参考匹配集合
target_series = table['Target word'].dropna()

# 标记目标词是否在匹配列表中
table['is_exact_match'] = target_series.isin(match_series)

# 提取第一组:精确匹配的条目
group1 = table[table['is_exact_match']].copy()

目标2的完整实现思路与代码

核心思路

  1. 筛选出未精确匹配的目标词
  2. 提前构建匹配列表的「首字符-对应条目」映射字典,避免重复遍历查询,提升效率
  3. 遍历每个未匹配的目标词,根据首字符查询映射字典,完成分组

完整代码

# 步骤1:筛选未精确匹配的目标词
unmatched_targets = table[~table['is_exact_match']]['Target word'].dropna().tolist()

# 步骤2:构建匹配列表的首字符映射字典
char_to_matches = {}
for word in match_series:
    first_char = word[0] if len(word) > 0 else ''
    if first_char not in char_to_matches:
        char_to_matches[first_char] = []
    char_to_matches[first_char].append(word)

# 步骤3:分组处理未匹配的目标词
group2 = []  # 存储格式:[(目标词, [匹配列表条目1, ...]), ...]
group3 = []

for target in unmatched_targets:
    if not target:  # 跳过空字符串
        continue
    first_char = target[0]
    if first_char in char_to_matches:
        group2.append( (target, char_to_matches[first_char]) )
    else:
        group3.append(target)

# 可选:将分组标签写入原表格,方便后续分析
table['group'] = '第三组'
table.loc[table['is_exact_match'], 'group'] = '第一组'

for target, _ in group2:
    table.loc[table['Target word'] == target, 'group'] = '第二组'

代码说明

  • 用char_to_matches字典预处理匹配列表,将查询时间复杂度从O(n)降到O(1)
  • 加入空值处理逻辑,避免空字符串引发的索引错误
  • 最后可选将分组标签合并回原DataFrame,便于后续的数据分析或可视化

内容的提问来源于stack exchange,提问作者Apples

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 12:30:41