如何在Python DataFrame中仅从转推提取首个提及用户,非转推留空
解决Pandas中仅提取转推内容的首个提及用户问题
需求说明
现有一个包含text_con_RT_t和rt列的Pandas DataFrame,rt列用于标记内容是否为转推(值为"RT"表示转推)。需要实现:
- 仅当
rt列为"RT"时,从text_con_RT_t中提取转推的首个提及用户,转换为小写并去除下划线等非字母字符; - 非转推的行,
usr_rt列留空。
示例DataFrame
import pandas as pd record = { 'text_con_RT_t' : ['RT @Blanc_: ramdon text #hashtag quiere ', '@GonM ramdon text', 'RT @IvEc: @GonzM ramdon text', 'hOLA ramdon text ' ], 'rt' : ['RT', '' , 'RT','' ] } dataframe2 = pd.DataFrame(record, columns = ['text_con_RT_t', 'rt'])
期望结果
| text_con_RT_t | rt | usr_rt |
|---|---|---|
| RT @Blanc_: ramdon text #hashtag quiere | RT | blanc |
| @GonM ramdon text | ||
| RT @IvEc: @GonzM ramdon text | RT | ivec |
| hOLA ramdon text |
错误结果(当前问题)
| text_con_RT_t | rt | usr_rt |
|---|---|---|
| RT @Blanc_: ramdon text #hashtag quiere | RT | blanc |
| @GonM ramdon text | gonm ramdon text | |
| RT @IvEc: @GonzM ramdon text | RT | ivec |
| hOLA ramdon text | NaN |
尝试过的代码
# 尝试1:错误的异常处理逻辑 try: dataframe2["usr_rt"] = dataframe2.text_con_RT_t.str.lower().str.split(':').str[0].str.split('@').str[1] except dataframe2["rt"]==None: dataframe2["usr_rt"] = "" # 尝试2:错误的整列条件判断 if dataframe2["rt"] == "RT": return (dataframe2["usr_rt"] == dataframe2.text_con_RT_t.str.split(':').str[0].str.split('@').str[1])
问题分析
- 缺少条件过滤:所有行都执行了提取操作,未区分是否为转推,导致非RT行错误提取内容;
- 不符合Pandas向量化逻辑:直接用
if整列判断、try-except处理DataFrame列的方式错误,Pandas需要用布尔索引、np.where或loc实现条件赋值; - 字符串分割不精准:
split(':')[0].split('@')[1]在无冒号的场景下会提取@后的所有内容,导致冗余数据。
解决方案
方案1:正则匹配 + np.where 条件赋值
使用正则精准匹配RT后的首个用户名,结合np.where实现条件处理:
import numpy as np import re # 定义提取并清理用户名的函数 def extract_rt_user(text): # 匹配RT后第一个@开头的用户名(到空格或冒号为止) match = re.search(r'RT @([^\s:]+)', text) if match: # 去除非字母字符并转小写 return re.sub(r'[^a-zA-Z]', '', match.group(1)).lower() return '' # 仅对rt为"RT"的行应用函数,其余行留空 dataframe2['usr_rt'] = np.where(dataframe2['rt'] == 'RT', dataframe2['text_con_RT_t'].apply(extract_rt_user), '')
方案2:布尔索引 + Pandas字符串方法
通过loc定位需要处理的行,用字符串方法提取并清理:
# 初始化usr_rt列为空字符串 dataframe2['usr_rt'] = '' # 生成转推行的布尔掩码 rt_mask = dataframe2['rt'] == 'RT' # 对转推行提取、清理用户名 dataframe2.loc[rt_mask, 'usr_rt'] = (dataframe2.loc[rt_mask, 'text_con_RT_t'] .str.extract(r'RT @([^\s:]+)', expand=False) .str.replace(r'[^a-zA-Z]', '', regex=True) .str.lower())
验证结果
执行上述任一方案后,均可得到符合预期的DataFrame,非转推行的usr_rt列保持为空,转推行精准提取并清理了首个提及用户。
内容的提问来源于stack exchange,提问作者Saríah
相关产品推荐
相关产品推荐

