复现维基百科印度邦名称抓取代码时遇AttributeError错误求助
解决网页抓取时的
AttributeError: ResultSet object has no attribute 'find_all'问题 嘿,我来帮你搞定这个错误!你遇到的这个问题其实是使用BeautifulSoup时非常常见的小坑,咱们一步步拆解解决。
问题背景
你尝试抓取维基百科上印度各邦及联邦属地的名称,但运行代码时触发了错误:
AttributeError: ResultSet object has no attribute 'find_all'
你的现有代码(部分):
# import library to query a website from urllib.request import urlopen # the url is stored in a variable called wiki wiki="https://en.wikipedia.org/wiki/List_of_state_and_union_territory_capitals_in_India"
错误原因分析
这个错误的核心是:你在一个ResultSet对象上调用了find_all()方法,但ResultSet是BeautifulSoup返回的「标签集合」(类似列表),它本身没有find_all方法——只有单个的Tag对象才有这个方法。
举个反例,如果你写了类似这样的代码就会触发错误:
# 错误示例:先得到ResultSet,再对它调用find_all() tables = soup.find_all('table', class_='wikitable') rows = tables.find_all('tr') # 这里tables是ResultSet,会报错!
修正后的完整代码
我帮你补全并修正了代码,结合BeautifulSoup(你之前可能漏导入了)来正确抓取数据:
from urllib.request import urlopen from bs4 import BeautifulSoup # 目标维基百科页面URL wiki = "https://en.wikipedia.org/wiki/List_of_state_and_union_territory_capitals_in_India" # 打开网页并解析HTML page = urlopen(wiki) soup = BeautifulSoup(page, 'html.parser') # 定位到包含目标数据的单个表格(用find()返回单个Tag,而非find_all()的集合) target_table = soup.find('table', class_='wikitable') # 从表格中获取所有行(此时target_table是单个Tag,调用find_all()是合法的) table_rows = target_table.find_all('tr') # 遍历行,提取邦/联邦属地名称(跳过表头行,取每行第一个单元格的文本) states_and_union_territories = [] for row in table_rows[1:]: # 从第2行开始(索引1),跳过表头 cells = row.find_all('td') if cells: # 过滤空行 territory_name = cells[0].text.strip() # 清理文本中的换行和空格 states_and_union_territories.append(territory_name) # 打印抓取到的结果 print("印度各邦及联邦属地名称:") for name in states_and_union_territories: print(name)
关键修正点
- 导入BeautifulSoup:你之前的代码没导入这个核心解析库,补上它才能正常解析HTML
- 用
find()替代find_all()获取父元素:先通过find()拿到单个表格Tag对象,再对它调用find_all()获取行,这才是正确的调用顺序 - 遍历ResultSet时操作单个Tag:对
table_rows(ResultSet)里的每个row(单个Tag)调用find_all('td'),这是合法且正确的用法
这样运行代码就能成功抓取到你需要的名称啦!
内容的提问来源于stack exchange,提问作者Ami
相关产品推荐
相关产品推荐

