You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理CSV数据拼接列并生成指定实体标注格式如何实现?

实现代码

首先确保你已经安装了pandas依赖,直接运行以下代码即可:

import pandas as pd

# 读取CSV文件,替换为你自己的文件路径
df = pd.read_csv("your_data.csv")

def row_process(row):
    # 转为字符串并去除首尾空白,避免多余空格影响下标计算
    house_content = str(row["house"]).strip()
    district_content = str(row["district"]).strip()
    
    # 拼接完整字符串,此处用单个空格分隔和示例规则对齐,可按需修改分隔符
    full_str = f"{house_content} {district_content}"
    
    # 计算两个字段的下标,默认遵循Python左闭右开切片规范
    house_start, house_end = 0, len(house_content)
    # 加1是两个字段中间的空格占位,修改分隔符的话这里对应调整偏移量
    district_start = house_end + 1
    district_end = district_start + len(district_content)
    
    # 绑定标签组装实体
    entities = [
        [(house_start, house_end), row["label1"]],
        [(district_start, district_end), row["label2"]]
    ]
    return (full_str, {"entities": entities})

# 逐行处理后转为列表格式,就是你要的最终输出
output = df.apply(row_process, axis=1).tolist()

# 测试打印验证
print(output)

注意事项

  • 如果house和district拼接不需要中间空格,把拼接语句的空格删掉,同时district_start改为house_end即可
  • 若你的标注要求是左闭右闭的下标格式,将house_end和district_end的计算结果各减1即可
  • 可以在函数内增加空值判断,跳过内容为空的字段对应的实体,避免生成无效标注

内容的提问来源于stack exchange,提问作者bellatrix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 14:45:04