如何抓取指定网站的美国各州数据表格?代码出现'NoneType'属性错误
问题原因
- 你使用的
jsx-a3119e4553b2cac7是前端框架生成的动态类名,每次页面加载都会随机变化,无法稳定定位表格元素。 - 目标网站存在基础反爬机制,直接用
requests.get请求会返回不含表格的页面,导致soup.find返回None。
修复方案
- 更换表格定位方式:使用页面中稳定存在的类名(比如
tp-table-body)或者标签结构来定位表格,避免依赖动态生成的类名。 - 添加请求头模拟浏览器:给
requests.get添加User-Agent等请求头,绕过基础反爬。
修改后代码
from bs4 import BeautifulSoup import requests import pandas as pd url = 'https://worldpopulationreview.com/states' # 添加请求头,模拟Chrome浏览器 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } page = requests.get(url, headers=headers) soup = BeautifulSoup(page.text, 'lxml') # 使用稳定的类名tp-table-body定位表格,去掉动态类名 table = soup.find('table', {'class': 'tp-table-body'}) if table is None: print("未找到表格,请检查请求头或页面结构是否变化") else: headers = [th.text.strip() for th in table.find_all('th')] df = pd.DataFrame(columns=headers) for row in table.find_all('tr')[1:]: row_data = [td.text.strip() for td in row.find_all('td')] df.loc[len(df)] = row_data print(df)
说明
- 新增的
headers模拟了真实浏览器的请求,避免被网站识别为爬虫。 - 表格定位改用
tp-table-body这个稳定类名,不会随页面渲染变化。 - 添加了
table是否为空的判断,避免直接调用find_all抛出异常。
内容的提问来源于stack exchange,提问作者user888469
相关产品推荐
相关产品推荐

