You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则移除DataFrame列星号空行并生成temp、word两列

DataFrame的localisation列处理方案

原代码出现大量None的原因

  • 没有边界判断:如果待处理的字符串没有换行符,调用split("\n")[1]会触发索引越界错误,无返回值得到None
  • 外层if逻辑冗余:re.findall返回的是匹配正则的子串列表,只要当前行包含正则外的特殊字符,就会导致if description.split("\n")[1] in substring判断不成立,函数无对应分支返回,默认得到None
  • 逻辑和需求不匹配:现有代码完全没有实现移除星号、过滤空行、匹配指定词汇的要求,逻辑写偏了

修正后完整代码

import pandas as pd
import re

# 给定匹配词列表
words = ['SECTION 11', 'CONE', 'BELLY', 'FIXED PLAN']

# 定义处理函数
def process_localisation(desc):
    # 处理temp列:取第一个换行符之后的所有内容,移除星号,过滤空行
    split_res = desc.split("\n", maxsplit=1) # 只分割1次,保留后续换行结构
    if len(split_res) < 2:
        temp = ""
    else:
        # 移除所有星号,按行拆分过滤空行后重新拼接
        temp_lines = [line.replace("*", "").strip() for line in split_res[1].split("\n")]
        temp = "\n".join([line for line in temp_lines if line])
    
    # 处理word列:匹配指定列表里的词,移除星号
    match_word = ""
    for w in words:
        if w in temp:
            match_word = w.replace("*", "")
            break # 匹配到第一个就终止,需要匹配所有可修改为列表收集
    return pd.Series([temp, match_word])

# 调用函数给DataFrame新增两列
df[["temp", "word"]] = df["localisation"].apply(process_localisation)

代码说明

  • 新增边界判断逻辑,完全避免索引越界问题,所有分支都有明确返回值,不会再出现None
  • 处理temp列时统一移除所有星号,逐行去除首尾空白后过滤空行,完全符合需求
  • 匹配word时直接遍历指定列表匹配,匹配结果自动做星号移除处理

内容的提问来源于stack exchange,提问作者grinim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 11:27:02