You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法抓取目标网站完整源码?添加请求头仍无效的技术求助

网页抓取问题:无法获取完整源码及解决方案

问题描述

无法抓取目标网站完整源码,打印response内容极短,无法获取有效信息。已添加User-Agent请求头但无效果,使用代码如下:

import requests
from bs4 import BeautifulSoup
import time

# User-Agent header
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.163 Safari/537.36'}

# Send an HTTP request to the webpage
response = requests.get('https://order.mikunisushi.com/menu/mikuni-folsom', headers=headers)

# Parse the HTML content of the webpage
soup = BeautifulSoup(response.text, 'html.parser')

# Find the product information element on the webpage
product_info = soup.find('div', class_='product__info')

if product_info:
  # Extract the product name and price from the element
  name = product_info.find('h1').text
  price = product_info.find('span', class_='price').text

  print(f'Product name: {name}')
  print(f'Product price: {price}')
else:
  print('Product information not found')

解决方案

这个网站的内容是JavaScript动态渲染的,requests只能获取初始静态HTML,无法加载JS生成的动态内容,因此需要用模拟真实浏览器的工具来抓取:

方法一:使用Selenium

Selenium可以模拟浏览器加载页面,等待JS执行完成后再获取完整DOM结构:

  1. 先安装依赖:
pip install selenium

同时需要下载对应浏览器的驱动(比如ChromeDriver),确保驱动版本与本地浏览器版本匹配。

  1. 改写后的代码示例:
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# 配置Chrome浏览器(启用无头模式,不弹出可视化窗口)
options = webdriver.ChromeOptions()
options.add_argument('--headless=new')
options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.163 Safari/537.36')

driver = webdriver.Chrome(options=options)

try:
    # 加载目标页面
    driver.get('https://order.mikunisushi.com/menu/mikuni-folsom')
    
    # 等待产品信息元素加载完成(最多等待10秒)
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, 'product__info'))
    )
    
    # 获取完整页面源码
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, 'html.parser')
    
    # 提取产品信息
    product_info = soup.find('div', class_='product__info')
    if product_info:
        name = product_info.find('h1').text
        price = product_info.find('span', class_='price').text
        print(f'Product name: {name}')
        print(f'Product price: {price}')
    else:
        print('Product information not found')
finally:
    # 关闭浏览器
    driver.quit()

注意事项

  • 动态页面抓取必须等待关键元素加载完成,避免因JS未执行完毕导致元素找不到
  • 可额外添加Referer、Accept-Language等请求头,进一步模拟真实用户请求
  • 避免短时间内频繁请求,可添加time.sleep()控制请求间隔,防止触发网站反爬机制

内容的提问来源于stack exchange,提问作者bakinglemoncookies

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 00:40:18