无法抓取目标网站完整源码?添加请求头仍无效的技术求助
网页抓取问题:无法获取完整源码及解决方案
问题描述
无法抓取目标网站完整源码,打印response内容极短,无法获取有效信息。已添加User-Agent请求头但无效果,使用代码如下:
import requests from bs4 import BeautifulSoup import time # User-Agent header headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.163 Safari/537.36'} # Send an HTTP request to the webpage response = requests.get('https://order.mikunisushi.com/menu/mikuni-folsom', headers=headers) # Parse the HTML content of the webpage soup = BeautifulSoup(response.text, 'html.parser') # Find the product information element on the webpage product_info = soup.find('div', class_='product__info') if product_info: # Extract the product name and price from the element name = product_info.find('h1').text price = product_info.find('span', class_='price').text print(f'Product name: {name}') print(f'Product price: {price}') else: print('Product information not found')
解决方案
这个网站的内容是JavaScript动态渲染的,requests只能获取初始静态HTML,无法加载JS生成的动态内容,因此需要用模拟真实浏览器的工具来抓取:
方法一:使用Selenium
Selenium可以模拟浏览器加载页面,等待JS执行完成后再获取完整DOM结构:
- 先安装依赖:
pip install selenium
同时需要下载对应浏览器的驱动(比如ChromeDriver),确保驱动版本与本地浏览器版本匹配。
- 改写后的代码示例:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # 配置Chrome浏览器(启用无头模式,不弹出可视化窗口) options = webdriver.ChromeOptions() options.add_argument('--headless=new') options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.163 Safari/537.36') driver = webdriver.Chrome(options=options) try: # 加载目标页面 driver.get('https://order.mikunisushi.com/menu/mikuni-folsom') # 等待产品信息元素加载完成(最多等待10秒) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'product__info')) ) # 获取完整页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 提取产品信息 product_info = soup.find('div', class_='product__info') if product_info: name = product_info.find('h1').text price = product_info.find('span', class_='price').text print(f'Product name: {name}') print(f'Product price: {price}') else: print('Product information not found') finally: # 关闭浏览器 driver.quit()
注意事项
- 动态页面抓取必须等待关键元素加载完成,避免因JS未执行完毕导致元素找不到
- 可额外添加
Referer、Accept-Language等请求头,进一步模拟真实用户请求 - 避免短时间内频繁请求,可添加
time.sleep()控制请求间隔,防止触发网站反爬机制
内容的提问来源于stack exchange,提问作者bakinglemoncookies
相关产品推荐
相关产品推荐

