You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Pandas高效提取HTML<a>标签中的文本内容?

方法一:正则表达式配合str.extract

你之前的正则(&gt;[A-Za-z])&lt;仅能匹配单个字母,无法覆盖空格、符号等内容,改用可匹配>到</a>之间所有非<字符的正则即可解决问题:

import pandas as pd

# 假设HTML数据存储在DataFrame的`html`列中
df = pd.DataFrame({
    'html': [
        '<a href="http://twitter.com/download/iphone" rel="nofollow">Twitter for iPhone</a>',
        '<a href="http://twitter.com" rel="nofollow">Twitter Web Client</a>',
        '<a href="http://vine.co" rel="nofollow">Vine - Make a Scene</a>',
        '<a href="https://about.twitter.com/products/tweetdeck" rel="nofollow">TweetDeck</a>'
    ]
})

# 用正则提取>和<之间的所有非<字符,分组捕获目标文本
df['text'] = df['html'].str.extract(r'>([^<]+)<')
print(df['text'])

输出结果:

0    Twitter for iPhone
1    Twitter Web Client
2    Vine - Make a Scene
3              TweetDeck
Name: text, dtype: object

方法二:使用BeautifulSoup(更稳定的HTML解析方式)

如果HTML结构存在变化(比如标签嵌套、属性顺序调整),正则容易失效,推荐用专门的HTML解析库BeautifulSoup,适配性更强:

先安装依赖库(未安装时执行):

pip install beautifulsoup4

编写解析代码:

import pandas as pd
from bs4 import BeautifulSoup

df = pd.DataFrame({
    'html': [
        '<a href="http://twitter.com/download/iphone" rel="nofollow">Twitter for iPhone</a>',
        '<a href="http://twitter.com" rel="nofollow">Twitter Web Client</a>',
        '<a href="http://vine.co" rel="nofollow">Vine - Make a Scene</a>',
        '<a href="https://about.twitter.com/products/tweetdeck" rel="nofollow">TweetDeck</a>'
    ]
})

# 定义函数提取a标签内的文本
def get_a_text(html):
    soup = BeautifulSoup(html, 'html.parser')
    a_tag = soup.find('a')
    return a_tag.get_text(strip=True) if a_tag else None

# 批量处理列数据
df['text'] = df['html'].apply(get_a_text)
print(df['text'])

输出结果与方法一一致,该方法能应对更复杂的HTML场景,避免正则表达式的局限性。

内容的提问来源于stack exchange,提问作者Khola

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 00:18:30