You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python移除字符串中URL及HTML标签无效,求解决方法

解决DataFrame中HTML文本清理问题

问题说明

我在DataFrame的某一列有多行带HTML转义字符的字符串,尝试用Python代码移除<p>、<a>标签、href属性及URL本身,但执行现有代码后内容毫无变化,请求解决:
原代码:

df["text"] = df["text"].str.replace(r'\s*https?://\S+(\s+|$)', ' ').str.strip()
df["text"] = df["text"].str.replace(r'\s*href=//\S+(\s+|$)', ' ').str.strip()

原始字符串示例:

&lt;p&gt;On 4 May 2019, &lt;a href=&quot;https://www.ft.com/content/26b6d24e-6d77-11e9-a9a5-351eeaef6d84&quot;&gt;The Financial Times&lt;/a&gt; (FT) reported that Huawei is planning to build a '400-person chip research and development factory' outside Cambridge. Planned to be operational by 2021, the factory will include an R&amp;amp;D centre and will be built on a 550-acre site reportedly purchased by Huawei in 2018 for &amp;pound;37.5 million. A Huawei spokesperson quoted in the FT article cited Huawei's long-term collaboration with Cambridge University, which includes a five-year, &amp;pound;25 million research partnership with BT, which launched a joint research group at the University of Cambridge. Read more about that partnership on this &lt;a href=&quot;https://chinatechmap.aspi.org.au/#/map/marker-1024&quot;&gt;map.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;In 2020 it was reported that the Huawei research and development center &lt;a href=&quot;https://archive.ph/wip/OrAd3&quot;&gt;received approval by a local council&lt;/a&gt; despite the nation&amp;rsquo;s ongoing security concerns around the Chinese company.&lt;/p&gt;
&lt;p&gt;Chinese state media later &lt;a href=&quot;https://web.archive.org/web/20190505143025/https://www.chinadaily.com.cn/a/201905/05/WS5ccecddfa3104842260b9df9.html&quot;&gt;reported that&lt;/a&gt; Huawei's expansion in Cambridge 'is part of a five-year, &amp;pound;3 billion investment plan for the UK that [Huawei] announced alongside [then] British Prime Minister Theresa May' in February 2018.&lt;/p&gt;

问题原因

  1. 原始字符串是HTML实体转义后的内容,比如&lt;p&gt;对应<p>、&quot;对应双引号,现有正则未处理这些转义字符,导致匹配失败。
  2. 现有正则的href=//\S+完全不符合实际格式,实际是href=&quot;URL&quot;,匹配规则错误。
  3. 直接用正则处理HTML标签容易遗漏嵌套标签、属性格式差异等边界情况。

解决方法

方法一:正则逐步清理(适合简单场景)

先反转义HTML实体,再针对性移除标签和URL:

import pandas as pd
import html

# 1. 反转义HTML实体,还原成正常HTML标签
df["text"] = df["text"].apply(html.unescape)

# 2. 移除<a>标签及其href属性,保留标签内的文本
df["text"] = df["text"].str.replace(r'<a\s+href="[^"]*">([^<]*)</a>', r'\1', regex=True)

# 3. 移除<p>标签
df["text"] = df["text"].str.replace(r'</?p>', '', regex=True)

# 4. 清理多余空格和换行
df["text"] = df["text"].str.strip().str.replace(r'\s+', ' ', regex=True)

方法二:用BeautifulSoup解析HTML(更可靠,推荐)

用专业HTML解析库处理,避免正则的局限性:

import pandas as pd
from bs4 import BeautifulSoup
import html

def clean_html(text):
    # 反转义HTML实体
    unescaped_text = html.unescape(text)
    # 解析HTML内容
    soup = BeautifulSoup(unescaped_text, 'html.parser')
    # 去掉<a>标签,保留里面的文本
    for a_tag in soup.find_all('a'):
        a_tag.unwrap()
    # 去掉<p>标签,保留里面的文本
    for p_tag in soup.find_all('p'):
        p_tag.unwrap()
    # 提取文本并清理多余空格
    cleaned_text = soup.get_text(strip=True)
    return ' '.join(cleaned_text.split())

# 应用到DataFrame列
df["text"] = df["text"].apply(clean_html)

效果验证

处理后示例文本会变成:

On 4 May 2019, The Financial Times (FT) reported that Huawei is planning to build a '400-person chip research and development factory' outside Cambridge. Planned to be operational by 2021, the factory will include an R&D centre and will be built on a 550-acre site reportedly purchased by Huawei in 2018 for £37.5 million. A Huawei spokesperson quoted in the FT article cited Huawei's long-term collaboration with Cambridge University, which includes a five-year, £25 million research partnership with BT, which launched a joint research group at the University of Cambridge. Read more about that partnership on this map. In 2020 it was reported that the Huawei research and development center received approval by a local council despite the nation’s ongoing security concerns around the Chinese company. Chinese state media later reported that Huawei's expansion in Cambridge 'is part of a five-year, £3 billion investment plan for the UK that [Huawei] announced alongside [then] British Prime Minister Theresa May' in February 2018.

内容的提问来源于stack exchange,提问作者George

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 19:03:36