使用Python BeautifulSoup下载指定JSON时获取到网页源码的问题
问题原因及解决方法
核心问题分析
- URL拼接错误:你拼接
json_url时重复了路径前缀,导致请求的是错误页面,返回网页源码而非JSON数据。比如锚点的href如果是/players/58392.json,你的拼接会生成https://slapshot.gg/players/58392/players/58392.json,这是无效地址。 - 动态元素未等待:即使使用Selenium,页面元素可能未完全加载就提取HTML,导致找不到JSON链接。
- 请求头缺失:直接用
requests.get请求时未携带浏览器标识,网站可能返回网页而非JSON。
修正后的代码
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import requests import time # Set up the WebDriver driver = webdriver.Firefox() try: # Open the webpage driver.get('https://slapshot.gg/players/58392') # 等待JSON链接加载完成(最多等10秒) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.LINK_TEXT, "JSON")) ) # 稍等确保元素稳定 time.sleep(1) # Get the HTML content html = driver.page_source # Parse the HTML using BeautifulSoup soup = BeautifulSoup(html, 'html.parser') # Find the anchor tag with the JSON string anchor_tag = soup.find("a", string="JSON") if anchor_tag: # 正确拼接URL json_href = anchor_tag["href"] json_url = f"https://slapshot.gg{json_href}" if json_href.startswith('/') else json_href # 模拟浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/115.0' } # Send a GET request to the JSON URL with headers json_response = requests.get(json_url, headers=headers) if json_response.status_code == 200: # 验证并保存JSON try: json_response.json() filename = json_url.split("/")[-1] with open(filename, "w", encoding="utf-8") as f: f.write(json_response.text) print("Downloaded the JSON file:", filename) except ValueError: print("Response is not valid JSON. Check the URL or headers.") else: print("Failed to retrieve JSON data. Status code:", json_response.status_code) else: print("JSON download link not found on the webpage.") finally: # Close the WebDriver driver.quit()
关键修正点
- 正确拼接URL:判断
href是否为相对路径,拼接根域名而非重复页面路径。 - 等待元素加载:使用
WebDriverWait确保JSON链接完全加载后再提取HTML。 - 添加请求头:模拟浏览器的
User-Agent,避免被网站拦截或返回非JSON内容。 - 验证JSON有效性:尝试解析响应内容,确认是否为合法JSON。
内容的提问来源于stack exchange,提问作者mboyle9310
相关产品推荐
相关产品推荐

