Python爬取Brawlhalla玩家数据异常问题求助
问题原因分析
你遇到的问题核心是目标网站采用JavaScript动态渲染数据:requests库只能获取页面的初始静态HTML,其中的玩家名称、等级都是页面加载前的占位数据,真实玩家信息是在页面加载完成后通过AJAX请求从后端接口获取并渲染到页面上的,直接解析静态自然会拿到错误的默认值。
解决方案
方案1:用Selenium模拟浏览器渲染(新手友好,直观)
Selenium可以模拟真实浏览器打开页面,等JavaScript渲染完成后再提取数据:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC code = input("enter the code :") url = f"https://brawlhallastats.herokuapp.com/player/?player={code}" # 初始化Chrome浏览器(需提前下载对应版本的chromedriver) driver = webdriver.Chrome() driver.get(url) try: # 等待元素加载完成,最多等10秒 playername = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "playerName")) ) playerLevel = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "playerLevel")) ) print("玩家名称:", playername.text.strip()) print("玩家等级:", playerLevel.text.replace("Level:", "").strip()) finally: driver.quit()
方案2:直接调用后端API(高效,推荐用于机器人开发)
通过浏览器开发者工具抓包,能发现页面会向https://brawlhallastats.herokuapp.com/api/player发送POST请求获取真实数据,直接调用这个接口比模拟浏览器更高效:
import requests code = input("enter the code :") api_url = "https://brawlhallastats.herokuapp.com/api/player" payload = {"player_id": code} headers = {"Content-Type": "application/json"} response = requests.post(api_url, json=payload) player_data = response.json() playername = player_data.get("name") playerLevel = player_data.get("level") print(f"玩家名称: {playername}") print(f"玩家等级: {playerLevel}")
测试玩家代码65340087时,该方案会直接返回正确的HetoskiWannaWeed和等级58。
补充说明
pandas无法识别表格也是同样原因:表格数据是JS动态生成的,静态HTML里没有完整的表格结构。- 方案2返回的JSON包含完整玩家数据,你可以按需提取胜率、英雄战绩等更多信息,更适配Discord机器人的功能开发需求。
内容的提问来源于stack exchange,提问作者Het
相关产品推荐
相关产品推荐

