You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Web Scraping问题:Python脚本无法获取浏览器Inspect显示的完整HTML

获取动态渲染页面的完整HTML

你的问题本质是目标网站采用JavaScript动态渲染内容:requests库只能获取服务器返回的初始静态HTML框架,而浏览器会自动执行页面中的JS脚本,加载并渲染出完整的页面内容,这就是你在Inspect中看到的完整DOM结构。

以下是三种可行的解决方案:

1. 用Selenium模拟真实浏览器

Selenium可以完全模拟浏览器的行为,执行JS并渲染出完整页面,适合快速解决问题。

步骤与代码示例:

  • 先安装依赖:pip install selenium
  • 下载对应浏览器的驱动(比如ChromeDriver,需与你的Chrome版本匹配)
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from bs4 import BeautifulSoup
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 初始化Chrome浏览器
service = Service('chromedriver.exe')  # 替换为你的驱动文件路径
driver = webdriver.Chrome(service=service)

# 访问目标页面
target_url = 'https://www.myauto.ge/ka/pr/89476234/iyideba-manqanebi-sedani-bmw-m5-2018-benzini-tbilisi?offerType=superVip'
driver.get(target_url)

# 显式等待页面关键元素加载完成(比固定sleep更可靠)
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, 'some-key-element'))  # 替换为页面中必有的元素类名/ID
    )
except:
    pass  # 超时后继续执行,避免报错

# 获取渲染后的完整HTML
full_html = driver.page_source

# 关闭浏览器
driver.quit()

# 用BeautifulSoup解析
soup = BeautifulSoup(full_html, 'html.parser')
# 后续即可正常提取所需数据

2. 使用Playwright(更现代的自动化工具)

Playwright是微软推出的自动化测试工具,API更简洁,内置完善的等待机制,支持多浏览器(Chrome、Firefox、Safari)。

步骤与代码示例:

  • 安装依赖:pip install playwright,然后执行playwright install安装浏览器驱动
from playwright.sync_api import sync_playwright
from bs4 import BeautifulSoup

target_url = 'https://www.myauto.ge/ka/pr/89476234/iyideba-manqanebi-sedani-bmw-m5-2018-benzini-tbilisi?offerType=superVip'

with sync_playwright() as p:
    # 启动浏览器(headless=True可后台运行)
    browser = p.chromium.launch(headless=False)
    page = browser.new_page()
    
    # 访问页面并等待网络空闲(确保内容加载完成)
    page.goto(target_url, wait_until='networkidle')
    
    # 获取完整HTML
    full_html = page.content()
    browser.close()

# 解析HTML
soup = BeautifulSoup(full_html, 'html.parser')

3. 直接调用网站API(最高效的方案)

很多动态网站会通过API接口加载数据,你可以跳过HTML渲染,直接请求API获取结构化数据(JSON格式),效率更高。

操作步骤:

  1. 打开浏览器DevTools(F12),切换到Network标签
  2. 刷新目标页面,筛选XHR/Fetch类型的请求
  3. 查找返回车辆详情数据的API接口(通常URL包含商品ID,比如89476234)
  4. 复制该API的请求URL和必要的请求头,用requests直接请求

代码示例:

import requests

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    # 可根据实际情况添加其他请求头,比如Referer、Cookie等
}

# 替换为你找到的API接口URL
api_url = 'https://api.myauto.ge/v1/items/89476234'
response = requests.get(api_url, headers=headers)

# 解析JSON数据
data = response.json()
# 直接提取所需字段,无需解析HTML

内容的提问来源于stack exchange,提问作者James

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 01:20:27