使用BeautifulSoup4与Python爬取网站时遭遇403 Forbidden错误的解决咨询
解决Pandas read_html触发的403 Forbidden错误
错误原因
你推测的完全正确:pd.read_html(URL)会独立发起一次新的HTTP请求,底层依赖urllib实现,而这次请求默认没有携带User-Agent,被网站反爬机制拦截,返回403错误。你之前用requests.get获取的页面内容并没有被read_html复用,这是核心问题。
解决方案一:复用已获取的HTML内容(推荐)
通过requests带User-Agent获取页面后,直接把HTML文本传给read_html,避免重复发起请求:
import requests from bs4 import BeautifulSoup import pandas as pd URL = "https://www.example.com/example" # 配置模拟浏览器的User-Agent headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 带请求头发起请求,检查请求状态 response = requests.get(URL, headers=headers) response.raise_for_status() # 请求失败时直接抛出异常 page_content = response.text # 保留你的BeautifulSoup操作(如果需要处理其他元素) soup = BeautifulSoup(page_content, "html.parser") tables = soup.find_all('table') target_table = soup.find('table', class_="table table-hover") # 传入已获取的HTML内容,而非URL df_list = pd.read_html( page_content, attrs={'class': 'table table-hover'}, flavor='bs4', thousands=',' ) # 输出结果 print(df_list[0].head())
解决方案二:给read_html的请求添加User-Agent(直接传URL场景)
如果必须直接通过URL调用read_html,可以自定义请求逻辑,确保请求携带User-Agent:
import pandas as pd import requests URL = "https://www.example.com/example" headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 用requests带请求头获取响应,再传给read_html response = requests.get(URL, headers=headers) df_list = pd.read_html( response.text, attrs={'class': 'table table-hover'}, flavor='bs4', thousands=',' ) print(df_list[0].head())
关键注意事项
- 务必给所有HTTP请求(包括
requests和read_html间接发起的)携带合法的User-Agent,模拟正常浏览器访问。 response.raise_for_status()可以快速定位请求失败的原因,避免后续代码在无效内容上执行。- 不要重复发起请求,既减轻服务器压力,也降低被反爬拦截的概率。
内容的提问来源于stack exchange,提问作者Jack
相关产品推荐
相关产品推荐

