如何用Python Pandas高效提取HTML<a>标签中的文本内容?
Pandas提取HTML标签内文本的解决方案
方法一:正则表达式配合str.extract
你之前的正则(>[A-Za-z])<仅能匹配单个字母,无法覆盖空格、符号等内容,改用可匹配>到</a>之间所有非<字符的正则即可解决问题:
import pandas as pd # 假设HTML数据存储在DataFrame的`html`列中 df = pd.DataFrame({ 'html': [ '<a href="http://twitter.com/download/iphone" rel="nofollow">Twitter for iPhone</a>', '<a href="http://twitter.com" rel="nofollow">Twitter Web Client</a>', '<a href="http://vine.co" rel="nofollow">Vine - Make a Scene</a>', '<a href="https://about.twitter.com/products/tweetdeck" rel="nofollow">TweetDeck</a>' ] }) # 用正则提取>和<之间的所有非<字符,分组捕获目标文本 df['text'] = df['html'].str.extract(r'>([^<]+)<') print(df['text'])
输出结果:
0 Twitter for iPhone 1 Twitter Web Client 2 Vine - Make a Scene 3 TweetDeck Name: text, dtype: object
方法二:使用BeautifulSoup(更稳定的HTML解析方式)
如果HTML结构存在变化(比如标签嵌套、属性顺序调整),正则容易失效,推荐用专门的HTML解析库BeautifulSoup,适配性更强:
先安装依赖库(未安装时执行):
pip install beautifulsoup4
编写解析代码:
import pandas as pd from bs4 import BeautifulSoup df = pd.DataFrame({ 'html': [ '<a href="http://twitter.com/download/iphone" rel="nofollow">Twitter for iPhone</a>', '<a href="http://twitter.com" rel="nofollow">Twitter Web Client</a>', '<a href="http://vine.co" rel="nofollow">Vine - Make a Scene</a>', '<a href="https://about.twitter.com/products/tweetdeck" rel="nofollow">TweetDeck</a>' ] }) # 定义函数提取a标签内的文本 def get_a_text(html): soup = BeautifulSoup(html, 'html.parser') a_tag = soup.find('a') return a_tag.get_text(strip=True) if a_tag else None # 批量处理列数据 df['text'] = df['html'].apply(get_a_text) print(df['text'])
输出结果与方法一一致,该方法能应对更复杂的HTML场景,避免正则表达式的局限性。
内容的提问来源于stack exchange,提问作者Khola
相关产品推荐
相关产品推荐

