Python网页爬取遇'NoneType'无'text'属性错误,如何获取动态ID标签文本?
解决Goodreads榜单爬取中动态ID投票数获取失败的问题
问题背景
爬取Goodreads的《纽约时报2015年虚构类畅销书榜单》时,尝试通过固定IDloading_link_849498获取投票数,运行代码后触发AttributeError: 'NoneType' object has no attribute 'text'错误——原因是每本书对应的投票链接ID为动态生成,无法用固定值匹配。
错误原因
原代码中使用固定ID查找投票链接:
vote = votes[0].find("a", {"id":"loading_link_849498"}).text.strip()
该ID仅对应第一本书的投票链接,其他书籍的投票链接ID是不同的动态值,导致后续循环中find()返回None,调用.text时触发错误。
解决方案
可以通过以下几种方式绕过动态ID的限制:
方法1:通过元素位置定位
投票链接是margin-top:5px的div下的第二个<a>标签(第一个是分数链接),直接通过索引获取:
vote_a = votes[0].findAll("a")[1] vote = vote_a.text.strip()
方法2:通过onclick属性特征匹配
投票链接的onclick属性包含list/list_book特征,利用BeautifulSoup的属性筛选功能:
vote_a = votes[0].find("a", onclick=lambda x: x and "list/list_book" in x) if vote_a: vote = vote_a.text.strip() else: vote = "无投票数据"
方法3:通过文本内容匹配
投票链接的文本包含people voted,直接匹配文本特征:
vote_a = votes[0].find("a", text=lambda x: x and "people voted" in x.strip()) if vote_a: vote = vote_a.text.strip() else: vote = "无投票数据"
修改后的完整代码
以方法1为例,修改后的代码如下:
import bs4 from urllib.request import urlopen as uReq from bs4 import BeautifulSoup as soup myurl = "https://www.goodreads.com/list/show/83612.NY_Times_Fiction_Best_Sellers_2015" uClient = uReq(myurl) page_html = uClient.read() uClient.close() page_soup = soup(page_html, "html.parser") tds = page_soup.findAll("td", {"width":"100%","valign":"top"}) for td in tds : # 获取书名 juduls = td.a.select("span") judul = juduls[0].text print(judul) # 获取作者 authors = td.find("div", class_="authorName__container").select("a") author = authors[0].text print(author) # 获取评分 rates = td.findAll("span", {"class":"greyText smallText uitext"}) rating = rates[0].text.strip().replace(",",".") print(rating) # 获取分数 scores = td.findAll("div", {"style":"margin-top: 5px"}) score = scores[0].span.a.text.strip().replace("score: ","").replace(",",".") print(score) # 获取投票数(修改后的部分) votes = td.findAll("div", {"style":"margin-top: 5px"}) vote_a = votes[0].findAll("a")[1] vote = vote_a.text.strip() print(vote) print("---")
内容的提问来源于stack exchange,提问作者Estri. P Lestari
相关产品推荐
相关产品推荐

