You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python DataFrame中仅从转推提取首个提及用户,非转推留空

解决Pandas中仅提取转推内容的首个提及用户问题

需求说明

现有一个包含text_con_RT_t和rt列的Pandas DataFrame,rt列用于标记内容是否为转推(值为"RT"表示转推)。需要实现:

  • 仅当rt列为"RT"时,从text_con_RT_t中提取转推的首个提及用户,转换为小写并去除下划线等非字母字符;
  • 非转推的行,usr_rt列留空。

示例DataFrame

import pandas as pd

record = { 
 'text_con_RT_t' : ['RT @Blanc_: ramdon text #hashtag quiere ', '@GonM ramdon text', 'RT @IvEc: @GonzM ramdon text', 'hOLA ramdon text ' ], 
 'rt' : ['RT', '' , 'RT','' ]
}
    
dataframe2 = pd.DataFrame(record, columns = ['text_con_RT_t', 'rt']) 

期望结果

text_con_RT_trtusr_rt
RT @Blanc_: ramdon text #hashtag quiereRTblanc
@GonM ramdon text
RT @IvEc: @GonzM ramdon textRTivec
hOLA ramdon text

错误结果(当前问题)

text_con_RT_trtusr_rt
RT @Blanc_: ramdon text #hashtag quiereRTblanc
@GonM ramdon textgonm ramdon text
RT @IvEc: @GonzM ramdon textRTivec
hOLA ramdon textNaN

尝试过的代码

# 尝试1:错误的异常处理逻辑
try:
   dataframe2["usr_rt"] = dataframe2.text_con_RT_t.str.lower().str.split(':').str[0].str.split('@').str[1]
except dataframe2["rt"]==None: 
   dataframe2["usr_rt"] = ""

# 尝试2:错误的整列条件判断
if dataframe2["rt"] == "RT":
  return (dataframe2["usr_rt"] == dataframe2.text_con_RT_t.str.split(':').str[0].str.split('@').str[1])

问题分析

  1. 缺少条件过滤:所有行都执行了提取操作,未区分是否为转推,导致非RT行错误提取内容;
  2. 不符合Pandas向量化逻辑:直接用if整列判断、try-except处理DataFrame列的方式错误,Pandas需要用布尔索引、np.where或loc实现条件赋值;
  3. 字符串分割不精准:split(':')[0].split('@')[1]在无冒号的场景下会提取@后的所有内容,导致冗余数据。

解决方案

方案1:正则匹配 + np.where 条件赋值

使用正则精准匹配RT后的首个用户名,结合np.where实现条件处理:

import numpy as np
import re

# 定义提取并清理用户名的函数
def extract_rt_user(text):
    # 匹配RT后第一个@开头的用户名(到空格或冒号为止)
    match = re.search(r'RT @([^\s:]+)', text)
    if match:
        # 去除非字母字符并转小写
        return re.sub(r'[^a-zA-Z]', '', match.group(1)).lower()
    return ''

# 仅对rt为"RT"的行应用函数,其余行留空
dataframe2['usr_rt'] = np.where(dataframe2['rt'] == 'RT',
                               dataframe2['text_con_RT_t'].apply(extract_rt_user),
                               '')

方案2:布尔索引 + Pandas字符串方法

通过loc定位需要处理的行,用字符串方法提取并清理:

# 初始化usr_rt列为空字符串
dataframe2['usr_rt'] = ''
# 生成转推行的布尔掩码
rt_mask = dataframe2['rt'] == 'RT'

# 对转推行提取、清理用户名
dataframe2.loc[rt_mask, 'usr_rt'] = (dataframe2.loc[rt_mask, 'text_con_RT_t']
                                     .str.extract(r'RT @([^\s:]+)', expand=False)
                                     .str.replace(r'[^a-zA-Z]', '', regex=True)
                                     .str.lower())

验证结果

执行上述任一方案后,均可得到符合预期的DataFrame,非转推行的usr_rt列保持为空,转推行精准提取并清理了首个提及用户。

内容的提问来源于stack exchange,提问作者Saríah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 05:20:43