You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从BrawlStats网站提取嵌套div中img的src链接失败,求解决

问题描述

目标:从指定个人主页提取特定蓝色排名logo的img src链接并存储到变量中,该图片的嵌套HTML结构如下:

<div class="mo25VS9slOfRz6jng3WTf">
    <img src="https://cdn.brawlstats.com/ranked-ranks/ranked_ranks_l_10.png" class="DPUFH-EhiGBBrkki4Gsaf">
    <div class="_3lMfMVxY-knKo2dnVHMCWG _21sSMvccqXG6cJU-5FNqzv" style="color:#FFFFFF;font-size:18px;">
    </div><!----></div>

尝试方案:使用requests+BeautifulSoup提取,但返回空结果,代码如下:

async def league_rank(interaction: discord.Interaction, tag: str):
    url = "https://brawlstats.com/profile/" + tag.upper()
    soup = BeautifulSoup(requests.get(url).content, "html.parser")
    all_imgs = [img["src"] for img in soup.select(".mo25VS9slOfRz6jng3WTf img")]
    print(all_imgs)
解决方法

问题核心是该网站依赖JavaScript动态渲染内容,requests只能获取初始静态HTML,此时目标元素还未加载,所以BeautifulSoup无法找到对应DOM节点。需要用能模拟浏览器渲染的工具获取完整页面,以下是两种可行方案:

方案1:使用Playwright(推荐)

Playwright是轻量的浏览器自动化工具,能直接获取JS渲染后的页面。先安装依赖:

pip install playwright
playwright install chromium

修改后的代码:

from playwright.async_api import async_playwright
from bs4 import BeautifulSoup

async def league_rank(interaction: discord.Interaction, tag: str):
    url = "https://brawlstats.com/profile/" + tag.upper()
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        await page.goto(url, wait_until="networkidle")  # 等待网络空闲确保内容加载完成
        html = await page.content()
        await browser.close()
        
        soup = BeautifulSoup(html, "html.parser")
        target_img = soup.select_one(".mo25VS9slOfRz6jng3WTf img")
        if target_img:
            img_src = target_img["src"]
            print(img_src)
            # 此处可将img_src存储到变量或执行后续逻辑
        else:
            print("未找到目标图片")

方案2:使用Selenium

Selenium是经典的浏览器自动化工具,安装依赖:

pip install selenium

需额外下载对应浏览器的驱动(如ChromeDriver),确保驱动与浏览器版本匹配。

修改后的代码:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
import asyncio

async def league_rank(interaction: discord.Interaction, tag: str):
    url = "https://brawlstats.com/profile/" + tag.upper()
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")  # 无头模式运行
    driver = webdriver.Chrome(options=chrome_options)
    driver.get(url)
    await asyncio.sleep(3)  # 等待元素加载,可根据实际情况调整时长
    html = driver.page_source
    driver.quit()
    
    soup = BeautifulSoup(html, "html.parser")
    target_img = soup.select_one(".mo25VS9slOfRz6jng3WTf img")
    if target_img:
        img_src = target_img["src"]
        print(img_src)
    else:
        print("未找到目标图片")
注意事项
  • 部分网站会检测自动化工具,可添加自定义User-Agent等配置模拟真实浏览器,避免被拦截。
  • 页面等待逻辑可根据实际加载速度调整,比如使用元素等待替代固定时长休眠,提升稳定性。

内容的提问来源于stack exchange,提问作者Zaid Hussain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 00:55:17