如何使用BeautifulSoup提取网页中多个H2标签的文本内容
问题解决方法
核心原因
find_all() 方法返回的是匹配到的所有元素组成的 ResultSet 列表对象,该对象本身没有 text 属性,必须遍历列表中的每个标签元素,单独提取文本内容。
修改后的完整代码
import requests from bs4 import BeautifulSoup import pandas as pd articlelist = [] url = 'https://www.angrybirds.com/blog/get-ready-angry-birds-movie-2-premiere-new-game-events/' r = requests.get(url) soup = BeautifulSoup(r.content, features='lxml') articles = soup.find_all('div', class_ = 'post-body__container') for item in articles: # 拿到所有h2元素 h2_elements = item.find_all('h2') # 遍历每个h2提取纯文本并去除首尾空白 h2_headings = [h2.text.strip() for h2 in h2_elements] # 按需求拼接为逗号分隔的字符串,和预期输出格式一致 h2_headings_str = ', '.join(h2_headings) article = { 'H2_Heading': h2_headings_str, } print('Added article:', article) articlelist.append(article) df = pd.DataFrame(articlelist) # 打印结果验证 print(df['H2_Heading'].iloc[0]) # df.to_csv('articlelist.csv', index=False) # print('Saved to csv')
对应问题说明
- 替换
find()为find_all()后遍历提取,解决只能获取第一个h2标题的问题 - 对每个h2元素单独调用
.text.strip(),得到纯文本内容,无多余HTML标签 - 避免直接对列表对象调用
text属性,解决AttributeError异常
运行代码后输出结果即为你需要的:ANGRY BIRDS 2, ANGRY BIRDS DREAM BLAST, ANGRY BIRDS FRIENDS, ANGRY BIRDS MATCH, ANGRY BIRDS BLAST, ANGRY BIRDS POP
内容的提问来源于stack exchange,提问作者MarkWP
相关产品推荐
相关产品推荐

