BeautifulSoup与Selenium无法抓取完整HTML及图片链接问题排查
问题原因
两种方案失效都是前端渲染机制导致的,没有特殊反爬玄学:
- 纯requests+BeautifulSoup拿不到内容:这个站是Next.js构建的客户端渲染站点,核心的玩家对局数据全是页面初始加载完成后,由浏览器发起异步JS请求拉取,再渲染到页面上的。requests只能拿到初始的空HTML架子,里面根本没有对局相关的DOM节点,选不到目标元素是正常结果。
- Selenium拿到base64占位图:站点用的是Next.js自带的
next/image图片组件,默认开启懒加载策略:不在当前视口范围内的图片,不会立刻加载真实地址,先塞一个1*1像素的透明base64图当占位符,等图片滚动到用户可见区域,才会把真实图片地址填充到srcset属性里。你之前的代码打开页面就立刻取源码,既没等对局列表加载完成,也没模拟滚动触发懒加载,拿到的自然全是占位图。
*额外说明:你判断站点没有公开API是误判,它只是没有对外提供开发者文档而已,前端拉取数据的接口是公开可访问的,直接调用接口拿结构化数据比爬DOM效率、稳定性都高很多。
实现方案
提供两个可直接跑通的方案,优先选第一个,不需要启动浏览器,资源消耗低、速度快。
方案1:直接调用站点内部数据接口(推荐)
打开Chrome开发者工具切到Network面板,筛选Fetch/XHR请求后刷新页面,就能找到站点拉取玩家数据的内部接口,直接请求该接口可以拿到结构化的JSON数据,完全不需要解析HTML。
实现代码:
import requests from urllib.parse import quote # 替换为需要查询的玩家名称即可 player_name = "ほばち" # 站点内部拉取玩家对局数据的接口 api_url = f"https://uniteapi.dev/api/player/{quote(player_name)}?region=jp" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36", "Referer": f"https://uniteapi.dev/p/{quote(player_name)}" } resp = requests.get(api_url, headers=headers) player_data = resp.json() # 最近50场对局数据在recent_matches字段中,直接统计即可 pokemon_count = {} for match in player_data.get("recent_matches", [])[:50]: use_pokemon = match.get("pokemon") if use_pokemon: pokemon_count[use_pokemon] = pokemon_count.get(use_pokemon, 0) + 1 print("最近50场宝可梦使用统计:") for pokemon, cnt in sorted(pokemon_count.items(), key=lambda x: x[1], reverse=True): print(f"{pokemon}: {cnt}场")
方案2:修正Selenium爬取逻辑
如果一定要通过解析渲染后DOM的方式获取数据,需要补充两个逻辑:一是加显式等待,等对局列表元素加载完成再取源码;二是模拟页面滚动,触发所有懒加载图片的真实地址加载。
修正后的代码:
from bs4 import BeautifulSoup as bs from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time url = "https://uniteapi.dev/p/%E3%81%BB%E3%81%B0%E3%81%A1" options = Options() options.add_argument('--headless=new') options.add_argument('--disable-gpu') options.add_argument('--window-size=1920,1080') driver = webdriver.Chrome(options=options) driver.get(url) # 等待首个宝可梦图片元素加载 try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'span img[alt="Played pokemon"]')) ) # 分段滚动页面,触发所有懒加载图片 for scroll_pos in range(0, 8000, 800): driver.execute_script(f"window.scrollTo(0, {scroll_pos})") time.sleep(0.3) # 滚动到页面底部等待1秒,确保所有资源加载完成 driver.execute_script("window.scrollTo(0, document.body.scrollHeight)") time.sleep(1) except Exception as e: print("页面加载异常:", e) page = driver.page_source driver.quit() soup = bs(page, 'html.parser') img_list = soup.select('span img[alt="Played pokemon"]') pokemon_count = {} for img in img_list[:50]: srcset = img.get("srcset", "") if "t_Square_" in srcset: # 从图片地址中提取宝可梦名称 pokemon_name = srcset.split("t_Square_")[1].split(".png")[0] pokemon_count[pokemon_name] = pokemon_count.get(pokemon_name, 0) + 1 print("最近50场宝可梦使用统计:") for pokemon, cnt in sorted(pokemon_count.items(), key=lambda x: x[1], reverse=True): print(f"{pokemon}: {cnt}场")
注意事项
- 爬取时控制请求频率,不要对站点造成不必要的压力
- 如果接口请求失败,先在浏览器开发者工具中核对接口路径、请求头参数,补全到代码中即可
- 长期使用优先选接口方案,不会因为前端页面结构调整导致解析逻辑失效
内容的提问来源于stack exchange,提问作者user19431042
相关产品推荐
相关产品推荐

