使用BeautifulSoup解析网页遇AttributeError: ResultSet无find_all属性求助
问题:提取class为name的p标签内a标签时触发AttributeError错误
尝试从class为name的<p>标签中提取内部<a>标签时,出现以下错误:
AttributeError : ResultSet object has no attribute 'find_all'
目标HTML代码片段
<div class="last_episodes loaddub"> <ul class="items"> <li> <div class="img"> <a href="/digimon-ghost-game-episode-36" title="Digimon Ghost Game"> <img alt="Digimon Ghost Game" src="https://gogocdn.net/cover/digimon-ghost-game.png"/> <div class="type ic-SUB"></div> </a> </div> <p class="name"><a href="/digimon-ghost-game-episode-36" title="Digimon Ghost Game">Digimon Ghost Game</a></p> <p class="episode">Episode 36</p> </li> <li> <div class="img"> <a href="/waccha-primagi-episode-41" title="Waccha PriMagi!"> <img alt="Waccha PriMagi!" src="https://gogocdn.net/cover/waccha-primagi.png"/> <div class="type ic-SUB"></div> </a> </div> <p class="name"><a href="/waccha-primagi-episode-41" title="Waccha PriMagi!">Waccha PriMagi!</a></p> <p class="episode">Episode 41</p> </li> <li> <div class="img"> <a href="/one-piece-episode-1027" title="One Piece"> <img alt="One Piece" src="https://gogocdn.net/images/anime/One-piece.jpg"/> <div class="type ic-SUB"></div> </a> </div> <p class="name"><a href="/one-piece-episode-1027" title="One Piece">One Piece</a></p> <p class="episode">Episode 1027</p> </li> <!-- 剩余HTML片段省略 --> </ul> </div>
初始代码
from urllib import response from venv import create from bs4 import BeautifulSoup import requests url = "https://gogoanime.gg?page=1" req = requests.get(url) Response = req.content soup = BeautifulSoup(Response, 'html.parser') p_tags = soup.find_all('p', class_='name') a_tags = p_tags.find_all('a') for link in a_tags: links = link.get('href') print(links)
尝试过的修改(未解决)
修改1:
p_tags = soup.find_all('p', class_='name') for a_tags in p_tags.find_all('a') print(a_tags)
修改2:能获取a标签但提取href报错,仅返回最后一个结果
p_tags = soup.find_all('p', class_='name') for a in p_tags: a_tags = a.find_all('a') print(len(a_tags))
解决方案
错误原因
soup.find_all('p', class_='name')返回的是ResultSet对象(类似列表的集合),它是多个Tag对象的容器,不能直接调用find_all()方法。必须遍历这个集合,对每个单独的<p>标签(Tag对象)执行查找操作。
正确代码1:遍历每个p标签提取a标签
from bs4 import BeautifulSoup import requests url = "https://gogoanime.gg?page=1" req = requests.get(url) soup = BeautifulSoup(req.content, 'html.parser') p_tags = soup.find_all('p', class_='name') for p in p_tags: # 每个p标签内只有一个a标签,用find()更高效 a_tag = p.find('a') if a_tag: href = a_tag.get('href') print(href)
正确代码2:用CSS选择器一步到位(更简洁)
直接通过CSS选择器定位到p.name内部的<a>标签,无需分步处理:
from bs4 import BeautifulSoup import requests url = "https://gogoanime.gg?page=1" req = requests.get(url) soup = BeautifulSoup(req.content, 'html.parser') # 直接选择所有class为name的p标签下的a标签 a_tags = soup.select('p.name a') for a in a_tags: href = a.get('href') print(href)
针对修改2的修复
修改2中a.find_all('a')返回的仍是ResultSet,需要从中取出单个a标签再提取href:
p_tags = soup.find_all('p', class_='name') for p in p_tags: a_tags = p.find_all('a') # 遍历每个a标签(虽然这里每个p只有一个) for a in a_tags: print(a.get('href'))
内容的提问来源于stack exchange,提问作者Omega500
相关产品推荐
相关产品推荐

