You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则表达式移除DataFrame文本列中HTTP及后续所有内容?

移除推文DataFrame中http及后续内容的正确正则方案

你的问题出在原正则表达式的匹配范围太窄,只覆盖了https加单个字符,所以没法彻底清除链接部分。要实现移除http(含https)及之后所有内容的需求,用下面的正则和代码即可:

archive_clean['text'] = archive_clean['text'].replace(r'\s*http.*$', '', regex=True)

正则说明:

  • \s*:匹配链接前的0个或多个空白字符,连带着把链接前的空格一起删掉,避免结果末尾留多余空格
  • http:匹配链接的起始标识(不管是http还是https,后面的s会被后续规则覆盖)
  • .*:匹配http之后的任意字符(除换行外)
  • $:锚定到字符串末尾,确保从http开始到结尾的所有内容都被替换为空

测试效果

用你给出的示例字符串测试:

原字符串:'This is Tilly. She's just checking pup on you. Hopes you're doing ok. If not, she's available for pats, snugs, boops, the whole bit. 13/10 https//(short url here i cannot post)'

处理后会得到:

'This is Tilly. She's just checking pup on you. Hopes you're doing ok. If not, she's available for pats, snugs, boops, the whole bit. 13/10'

内容的提问来源于stack exchange,提问作者Raghad Salem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 00:05:23