You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

非结构化数据处理与清洗:Pandas中混杂标记文本的格式化问题

你可以使用Python的pandas、re和html库完成文本清洗和对齐输出,完整实现代码如下:

import pandas as pd
import re
import html

def clean_music_text(raw_text_series):
    # 合并所有原始文本,按序号拆分单条记录
    full_raw_text = '\n'.join(raw_text_series.tolist())
    record_pattern = re.compile(r'^\s*(\d+)\.\s+(.*?)(?=\s*\d+\.\s+|\Z)', re.DOTALL | re.MULTILINE)
    raw_records = record_pattern.findall(full_raw_text)

    processed = []
    for seq, content in raw_records:
        # 解析HTML转义字符
        content = html.unescape(content)
        # 将HTML标签替换为专属分隔符,精准拆分创作者/作品字段
        content = re.sub(r'<[^>]+>', '|||', content, flags=re.DOTALL)
        # 过滤空值拆分字段
        parts = [p.strip() for p in content.split('|||') if p.strip()]
        # 清洗创作者字段
        creator = re.sub(r'^[*>"\s]+', '', parts[0]).replace(' / ', ', ').strip('" ')
        # 清洗作品字段
        work = ' '.join(parts[1:]).strip('" ') if len(parts) > 1 else ''
        processed.append({
            '序号': int(seq),
            '创作者': creator,
            '作品信息': work
        })
    
    # 转为DataFrame方便后续处理
    clean_df = pd.DataFrame(processed)
    # 计算对齐宽度,可根据需求调整偏移量
    max_creator_width = clean_df['创作者'].str.len().max() + 4
    # 格式化输出
    output_lines = []
    for _, row in clean_df.iterrows():
        line = f"{row['序号']}.  {row['创作者']:<{max_creator_width}}{row['作品信息']}"
        output_lines.append(line)
        print(line)
    return clean_df, output_lines

# 调用示例:假设你的原始数据存在df的raw列中
# clean_df, output = clean_music_text(df['raw'])

核心逻辑说明:

  • 先按序号规则合并同一条的多行内容,解决原始数据断行问题
  • 替换HTML标签为特殊分隔符,避免直接删除标签后无法拆分创作者和作品信息
  • 统一清洗转义字符、多余特殊符号、冗余空格后,按最大创作者字段长度左对齐输出
    如果你的实际数据存在更多特殊格式,微调对应正则规则即可适配。

内容的提问来源于stack exchange,提问作者matheus james

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 06:09:03