如何从Pandas列中提取特定格式的文本内容?
Pandas列数据解析方案
第一种结果提取
要得到不带末尾用户标识的内容,直接按 - <@ 分割字符串取前半部分即可,代码示例:
import pandas as pd # 模拟目标数据 df = pd.DataFrame({ 'col': ['**<https://stackoverflow.com/> Stackoverflow runs the world --website --q 3** - <@123456789> (today)'] }) # 提取第一种结果 df['result1'] = df['col'].str.split(r' - <@', expand=True)[0].str.strip('*') print(df['result1'].iloc[0])
输出结果:
<https://stackoverflow.com/> Stackoverflow runs the world --website --q 3
理想结果提取
需要同时去除链接的尖括号、以及所有--开头的后缀,步骤如下:
- 先分割掉末尾用户标识并移除首尾星号
- 用正则匹配出链接和
--之前的文本内容
代码示例:
import re # 定义提取函数 def extract_ideal(text): # 第一步:清理末尾标识和首尾星号 cleaned = text.split(r' - <@')[0].strip('*') # 第二步:匹配目标内容 match = re.match(r'<(https?://[^>]*)>(.*?)(?= --|$)', cleaned) if match: return f"{match.group(1)} {match.group(2).strip()}" return cleaned df['result_ideal'] = df['col'].apply(extract_ideal) print(df['result_ideal'].iloc[0])
输出结果:
https://stackoverflow.com/ Stackoverflow runs the world
该正则可适配单个或多个--后缀的情况,只要--是后缀起始标识就能正确截取。
内容的提问来源于stack exchange,提问作者rafaelHTML
相关产品推荐
相关产品推荐

