使用BeautifulSoup爬取Investing.com时tables[0]报404,tables[1]正常
解决Investing.com巴西CDS数据爬取的404与表格定位问题
1. 404错误的核心排查点
- 地域IP限制:Investing.com会根据访问IP的地域返回差异化内容,你的IP可能被定向到了无1年期CDS历史数据的页面,而他人IP能正常访问。直接在浏览器打开目标URL(如
https://br.investing.com/rates-bonds/brazil-cds-1-year-usd-historical-data),验证是否真的返回404。 - UA头适配问题:你当前使用的是Linux环境的Chrome UA,Windows环境下建议换成对应系统的UA,比如
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36,服务器可能对不同系统的UA做了内容分流。
2. 表格定位的可靠优化方案
不要用索引[0]硬定位表格(页面结构会因地域、会话变化),改用类名+字段匹配的精准方式:
# 替换原table查找逻辑 table = soup.find("table", class_="datatable_table__D_jso") # 类名可通过浏览器F12查看当前页面实际值 if not table: # 兼容不同页面结构,遍历找包含目标字段的表格 for tbl in soup.find_all("table"): if "Último" in str(tbl) and "Data" in str(tbl): table = tbl break
这种方式能确保找到包含你需要的Último和Data字段的表格,而非依赖固定索引顺序。
3. 环境相关的额外处理
- 添加会话Cookie支持:Investing.com可能需要会话Cookie才能返回正确内容,改用
requests库(比urllib更易用,支持会话保持):
import requests from bs4 import BeautifulSoup import pandas as pd from io import StringIO lista_cds = ['cds-1-year', 'cds-2-year', 'cds-3-year', 'cds-4-year', 'cds-5-year', 'cds-7-year', 'cds-10-year'] headers = {'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'} session = requests.Session() lista_dfs = [] for ano_cds in lista_cds: url = f'https://br.investing.com/rates-bonds/brazil-{ano_cds}-usd-historical-data' # 先访问主页获取会话Cookie session.get('https://br.investing.com/', headers=headers) response = session.get(url, headers=headers) response.raise_for_status() # 主动抛出HTTP错误,便于调试 soup = BeautifulSoup(response.text, features='lxml') # 精准查找目标表格 table = soup.find("table", class_="datatable_table__D_jso") if not table: for tbl in soup.find_all("table"): if "Último" in str(tbl): table = tbl break if table: df_cds = pd.read_html(StringIO(str(table)))[0][['Último', 'Data']] lista_dfs.append(df_cds) else: print(f"未找到{ano_cds}的目标表格")
- 检查网络代理:如果你的网络使用了代理,可能导致服务器返回异常内容,关闭代理后重试。
内容的提问来源于stack exchange,提问作者jaokz
相关产品推荐
相关产品推荐

