You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中合并explode后的行:提取Twitter提及用户名至原行

没问题,我帮你搞定这个提取username又不拆分行的需求!

解决方案:提取提及用户并保留原DataFrame行数

你遇到的痛点很典型——explode虽然能拆分嵌套列表,但会破坏原数据的行结构。我们可以用apply结合列表推导式,直接在原行内提取并合并所有username,完全保留原DataFrame的行数。

方法1:提取为列表形式(推荐)

这种方式会把同一行的所有username存成一个列表,方便后续的数据分析操作(比如统计提及频次):

import pandas as pd

# 假设你的DataFrame名为df,目标列是'Mentioned Users'
df['Mentioned Usernames'] = df['Mentioned Users'].apply(
    # 逐行处理:如果是None返回空列表,否则遍历字典提取username
    lambda x: [user['username'] for user in x] if x is not None else []
)

方法2:提取为逗号分隔的字符串

如果需要更直观的文本格式(比如导出成报表),可以把列表转成逗号分隔的字符串:

df['Mentioned Usernames'] = df['Mentioned Users'].apply(
    lambda x: ', '.join([user['username'] for user in x]) if x is not None else ''
)

额外优化:处理异常情况

如果你的数据里存在没有username键的字典(避免触发KeyError),可以用get方法设置默认值:

df['Mentioned Usernames'] = df['Mentioned Users'].apply(
    lambda x: [user.get('username', '未知用户') for user in x] if x is not None else []
)

示例效果

假设原数据是这样:

TweetMentioned Users
Hello @user1 @user2[{'username': 'user1'}, {'username': 'user2'}]
Hi thereNone
Hey @user3[{'username': 'user3'}]

处理后会得到:

TweetMentioned UsersMentioned Usernames
Hello @user1 @user2[{'username': 'user1'}, {'username': 'user2'}]['user1', 'user2']
Hi thereNone[]
Hey @user3[{'username': 'user3'}]['user3']

(如果用字符串格式,第三列会变成user1, user2、''、user3)

这种方式既保证了原数据的行数不变,又完美提取了所有提及的username,完全适配你的20000条Twitter数据场景。

内容的提问来源于stack exchange,提问作者user14913431

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:15:17