如何用正则表达式移除DataFrame文本列中HTTP及后续所有内容?
移除推文DataFrame中http及后续内容的正确正则方案
你的问题出在原正则表达式的匹配范围太窄,只覆盖了https加单个字符,所以没法彻底清除链接部分。要实现移除http(含https)及之后所有内容的需求,用下面的正则和代码即可:
archive_clean['text'] = archive_clean['text'].replace(r'\s*http.*$', '', regex=True)
正则说明:
\s*:匹配链接前的0个或多个空白字符,连带着把链接前的空格一起删掉,避免结果末尾留多余空格http:匹配链接的起始标识(不管是http还是https,后面的s会被后续规则覆盖).*:匹配http之后的任意字符(除换行外)$:锚定到字符串末尾,确保从http开始到结尾的所有内容都被替换为空
测试效果
用你给出的示例字符串测试:
原字符串:'This is Tilly. She's just checking pup on you. Hopes you're doing ok. If not, she's available for pats, snugs, boops, the whole bit. 13/10 https//(short url here i cannot post)'
处理后会得到:
'This is Tilly. She's just checking pup on you. Hopes you're doing ok. If not, she's available for pats, snugs, boops, the whole bit. 13/10'
内容的提问来源于stack exchange,提问作者Raghad Salem
相关产品推荐
相关产品推荐

