能否抓取特定网页?用BeautifulSoup抓城市数据遇IndexError问题求助
解决Web Scraping中的IndexError问题
背景情况
我想要自动抓取并保存http://www.dataforcities.org/网站里的城市数据,于是用BeautifulSoup库尝试获取http://open.dataforcities.org/details?4[]=2016页面的数据,初始代码是这样的:
import urllib2 from BeautifulSoup import BeautifulSoup soup = BeautifulSoup(urllib2.urlopen('http://open.dataforcities.org/details?4[]=2016').read())
后来参考《Web scraping with Python》里的示例写了下面的代码,结果直接触发了IndexError:
soup = BeautifulSoup(urllib2.urlopen('http://open.dataforcities.org/details?4[]=2016').read()) for row in soup('table', {'class': 'metrics'})[0].tbody('tr'): tds = row('td') print tds[0].string, tds[1].string
对应的错误信息如下:
IndexError Traceback (most recent call last) <ipython-input-71-d688ff354182> in <module>() ----> 1 for row in soup('table', {'class': 'metrics'})[0].tbody('tr'): 2 tds = row('td') 3 print tds[0].string, tds[1].string IndexError: list index out of range
问题原因
这个错误的核心很明确:soup('table', {'class': 'metrics'})返回的是空列表,你直接用[0]去取第一个元素,自然会触发索引越界错误。通常有两种可能性:
- 目标页面里根本不存在class为
metrics的table元素——可能是你记错了类名,或者页面结构已经和书中示例不一样了; - 页面内容是通过JavaScript动态渲染的,
urllib2只能获取初始HTML源码,拿不到JS加载后的实际内容。
解决方案步骤
1. 先确认页面的实际结构
打开目标页面,按下F12调出浏览器开发者工具,切换到「Elements」标签页,搜索是否有class为metrics的table。如果找不到,就需要重新定位数据所在容器——比如检查table的其他属性(id、其他class),或者确认数据是不是在其他标签里。
2. 增加代码的容错性
就算找到了正确元素,直接用[0]取列表元素也很危险,建议先判断列表是否为空,再进行后续操作,修改后的代码如下:
import urllib2 from BeautifulSoup import BeautifulSoup soup = BeautifulSoup(urllib2.urlopen('http://open.dataforcities.org/details?4[]=2016').read()) # 先获取所有符合条件的table target_tables = soup.find_all('table', {'class': 'metrics'}) if target_tables: # 取第一个table table = target_tables[0] # 获取tbody里的所有行 rows = table.tbody.find_all('tr') for row in rows: tds = row.find_all('td') # 确保每行至少有2个td再打印 if len(tds) >= 2: print tds[0].string, tds[1].string else: print("未找到class为metrics的table,请检查页面结构")
3. 处理动态加载的情况
如果检查后发现页面确实有这个table,但urllib2获取的源码里没有,那就是动态加载的问题。这种情况下,你可以改用Selenium模拟浏览器渲染页面,或者用requests-html库(支持JS渲染)。举个Selenium的简单示例:
from selenium import webdriver from BeautifulSoup import BeautifulSoup # 初始化浏览器(需要提前下载对应浏览器的driver,比如ChromeDriver) driver = webdriver.Chrome() driver.get('http://open.dataforcities.org/details?4[]=2016') # 获取渲染后的页面源码 soup = BeautifulSoup(driver.page_source, 'html.parser') # 后续操作和之前一样 target_tables = soup.find_all('table', {'class': 'metrics'}) # ... 省略后续处理代码 # 记得关闭浏览器 driver.quit()
内容的提问来源于stack exchange,提问作者emax
相关产品推荐
相关产品推荐

