You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将分词后的文本列关联回原始数据表

解决方案

用Pandas快速实现分词与Code关联展开

针对你的需求,最简便的方式是利用Pandas的explode()方法,直接将分词后的列表拆分为多行,同时保留原始的Code关联关系。以下是完整实现步骤:

  1. 准备环境与数据
    先确保安装了Pandas,再构造或读取原始数据表:

    import pandas as pd
    
    # 示例原始数据
    data = {
        'Code': ['ST-441', 'St-432'],
        'Text': ['Purpose of your visit mentioned', 'Describe how and where it happened']
    }
    df = pd.DataFrame(data)
    
  2. 修正分词函数并生成分词列
    你的原分词函数存在参数引用错误,修正后将每个文本行拆分为词列表,新增一列存储分词结果:

    def clean_text(text):
        # 可在此加入自定义清洗逻辑(如去停用词、大小写统一等)
        return text.split(' ')
    
    # 生成分词列
    df['Tokens'] = df['Text'].apply(clean_text)
    
  3. 展开分词为多行并关联Code
    使用explode()方法将分词列表拆分为独立行,同时保留对应Code:

    # 展开分词列,并重命名为Text得到目标格式
    result_df = df.explode('Tokens').rename(columns={'Tokens': 'Text'})
    # 可选:重置索引
    result_df = result_df.reset_index(drop=True)
    # 可选:统一Code大小写(如全部转大写)
    result_df['Code'] = result_df['Code'].str.upper()
    
  4. 查看结果
    打印结果即可得到你需要的格式:

    print(result_df[['Code', 'Text']])
    

    输出示例:

    Code       Text
    0  ST-441    Purpose
    1  ST-441         of
    2  ST-441        your
    3  ST-441       visit
    4  ST-441  mentioned
    5  ST-432    Describe
    6  ST-432         how
    7  ST-432         and
    8  ST-432        where
    9  ST-432          it
    10 ST-432    happened
    

若已提前得到分词列表

如果已经有独立的Code列表和分词列表,可直接构造DataFrame:

# 假设你已准备好这两个列表
codes = ['ST-441']*4 + ['ST-432']*6
tokens = ['Purpose', 'of', 'your', 'visit', 'mentioned', 'Describe', 'how', 'and', 'where', 'it', 'happened']

result_df = pd.DataFrame({'Code': codes, 'Text': tokens})

内容的提问来源于stack exchange,提问作者Dhanya_mj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 03:01:06