You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于匹配列的子串提取与多标签标记Python代码求助

Pandas多匹配标签与文本提取解决方案

需求说明

  1. 依据species_id列匹配指定ID子串列表,为specie列添加对应标签:匹配animal列表ID标记animal_customer,匹配human列表ID标记human_customer;同一行匹配多个ID时,标签以列表形式展示;无匹配时标记unknown。
  2. 匹配成功时,从story列的长字符串中提取匹配ID左侧1个、右侧3个空格分隔的子串存入extract列;无匹配时填充I.D. not found。
  3. 示例匹配列表:
    • animal = ["0325", "9985"]
    • human = ["9984", "1859"]

原始表示例

species_idstory
0325;9984The quick 0325 brown fox jumps over the 9984 lazy dog and runs far away
9985A small 9985 cat sits on the windowsill watching birds fly by
1234No matching IDs in this story at all

目标结果表示例

species_idstoryspecieextract
0325;9984The quick 0325 brown fox jumps over the 9984 lazy dog and runs far away["animal_customer", "human_customer"]["quick 0325 brown fox jumps", "the 9984 lazy dog and"]
9985A small 9985 cat sits on the windowsill watching birds fly by["animal_customer"]["small 9985 cat sits on"]
1234No matching IDs in this story at all"unknown""I.D. not found"

现有代码问题分析

  • 通过筛选子DataFrame再拼接的方式,会导致匹配多个标签的行重复出现
  • 未处理无匹配的行,这类行直接被排除在结果之外
  • 无法实现同一行多标签的列表展示形式

修正后代码

import pandas as pd
import re

# 定义匹配ID与对应标签的映射
animal_ids = ["0325", "9985"]
human_ids = ["9984", "1859"]
id_label_map = {}
for id_str in animal_ids:
    id_label_map[id_str] = "animal_customer"
for id_str in human_ids:
    id_label_map[id_str] = "human_customer"

# 构建示例原始数据
raw_data = {
    "species_id": ["0325;9984", "9985", "1234"],
    "story": [
        "The quick 0325 brown fox jumps over the 9984 lazy dog and runs far away",
        "A small 9985 cat sits on the windowsill watching birds fly by",
        "No matching IDs in this story at all"
    ]
}
df = pd.DataFrame(raw_data)

# 定义逐行处理函数
def process_single_row(row):
    # 拆分当前行的所有species_id
    split_ids = row["species_id"].split(";")
    matched_labels = []
    extracted_segments = []
    
    for target_id in split_ids:
        if target_id in id_label_map:
            # 记录匹配到的标签
            matched_labels.append(id_label_map[target_id])
            # 正则匹配目标ID左1右3个空格分隔的子串
            regex_pattern = re.compile(r'(\w+) ' + re.escape(target_id) + r' (\w+) (\w+) (\w+)')
            match_result = regex_pattern.search(row["story"])
            if match_result:
                # 拼接提取到的内容
                extracted = f"{match_result.group(1)} {target_id} {match_result.group(2)} {match_result.group(3)} {match_result.group(4)}"
                extracted_segments.append(extracted)
    
    # 处理最终标签:去重后转列表,无匹配则为unknown
    final_specie = list(set(matched_labels)) if matched_labels else "unknown"
    # 处理最终提取内容:无匹配则填充指定文本
    final_extract = extracted_segments if extracted_segments else "I.D. not found"
    
    return pd.Series([final_specie, final_extract], index=["specie", "extract"])

# 应用函数到每行,生成新列
df[["specie", "extract"]] = df.apply(process_single_row, axis=1)

# 打印结果
print(df)

代码关键说明

  1. 用字典映射ID与标签,避免重复判断,提升效率
  2. 采用apply逐行处理,从根源避免行重复问题
  3. 正则表达式精准定位目标ID前后的空格分隔词,确保提取内容符合要求
  4. 统一处理无匹配场景,填充对应默认值
  5. 对多匹配标签去重,保证列表内容唯一

内容的提问来源于stack exchange,提问作者Random Person

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 13:07:44