如何使用BeautifulSoup正确提取find_all返回所有元素的.text内容
错误原因
BeautifulSoup的find_all()方法返回的是ResultSet对象,本质是存储了所有匹配DOM元素的列表容器,只有列表内的单个Tag元素才支持调用.text属性提取文本,直接对整个列表调用.text就会触发属性不存在的报错。
解决方法
你只需要遍历find_all()返回的所有元素,逐个提取文本即可,根据使用场景可以选择下面两种常用写法:
1. 将所有文本存入列表供后续使用
import bs4 from bs4 import BeautifulSoup import requests import lxml vegas_insider = requests.get('https://www.vegasinsider.com/nfl/matchups/', 'r').text soup = BeautifulSoup(vegas_insider, 'lxml') # 先获取所有匹配的td元素集合 matchup_eles = soup.find_all('td', class_ = 'viHeaderNorm') # 遍历每个元素提取文本,strip()用于去除首尾空白字符,不需要可删除 matchup_texts = [ele.text.strip() for ele in matchup_eles] # 直接打印整个文本列表 print(matchup_texts)
2. 逐个打印所有匹配元素的文本
matchup_eles = soup.find_all('td', class_ = 'viHeaderNorm') for ele in matchup_eles: print(ele.text.strip())
内容的提问来源于stack exchange,提问作者RookiePython
相关产品推荐
相关产品推荐

