You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python BeautifulSoup下载指定JSON时获取到网页源码的问题

问题原因及解决方法

核心问题分析

  • URL拼接错误:你拼接json_url时重复了路径前缀,导致请求的是错误页面,返回网页源码而非JSON数据。比如锚点的href如果是/players/58392.json,你的拼接会生成https://slapshot.gg/players/58392/players/58392.json,这是无效地址。
  • 动态元素未等待:即使使用Selenium,页面元素可能未完全加载就提取HTML,导致找不到JSON链接。
  • 请求头缺失:直接用requests.get请求时未携带浏览器标识,网站可能返回网页而非JSON。

修正后的代码

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import requests
import time

# Set up the WebDriver
driver = webdriver.Firefox()

try:
    # Open the webpage
    driver.get('https://slapshot.gg/players/58392')
    
    # 等待JSON链接加载完成(最多等10秒)
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.LINK_TEXT, "JSON"))
    )
    
    # 稍等确保元素稳定
    time.sleep(1)
    
    # Get the HTML content
    html = driver.page_source

    # Parse the HTML using BeautifulSoup
    soup = BeautifulSoup(html, 'html.parser')

    # Find the anchor tag with the JSON string
    anchor_tag = soup.find("a", string="JSON")

    if anchor_tag:
        # 正确拼接URL
        json_href = anchor_tag["href"]
        json_url = f"https://slapshot.gg{json_href}" if json_href.startswith('/') else json_href

        # 模拟浏览器请求头
        headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/115.0'
        }
        
        # Send a GET request to the JSON URL with headers
        json_response = requests.get(json_url, headers=headers)

        if json_response.status_code == 200:
            # 验证并保存JSON
            try:
                json_response.json()
                filename = json_url.split("/")[-1]
                with open(filename, "w", encoding="utf-8") as f:
                    f.write(json_response.text)
                print("Downloaded the JSON file:", filename)
            except ValueError:
                print("Response is not valid JSON. Check the URL or headers.")
        else:
            print("Failed to retrieve JSON data. Status code:", json_response.status_code)
    else:
        print("JSON download link not found on the webpage.")
finally:
    # Close the WebDriver
    driver.quit()

关键修正点

  • 正确拼接URL:判断href是否为相对路径,拼接根域名而非重复页面路径。
  • 等待元素加载:使用WebDriverWait确保JSON链接完全加载后再提取HTML。
  • 添加请求头:模拟浏览器的User-Agent,避免被网站拦截或返回非JSON内容。
  • 验证JSON有效性:尝试解析响应内容,确认是否为合法JSON。

内容的提问来源于stack exchange,提问作者mboyle9310

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 13:30:14