为何无法抓取网站搜索结果页的展示结果URL?
解决方案:抓取sobrico搜索结果页的首个产品URL
问题根源
- 你使用的URL包含
#,这部分属于前端锚点参数,不会被发送到服务器,服务器返回的只是网站初始静态页面,不包含搜索结果内容。 - 搜索结果是通过JavaScript动态渲染加载的,
requests仅能获取静态HTML,无法执行JS生成页面内容,因此BeautifulSoup无法解析到搜索结果。
可行解决方法
方法1:直接调用搜索API(推荐)
通过浏览器开发者工具的网络面板,可找到网站搜索时调用的API接口,直接请求该接口获取结构化的搜索结果,无需处理JS渲染。
示例代码:
import requests search_query = "2608664131" # 网站实际搜索API接口 api_url = f"https://www.sobrico.com/search/ajax?query={search_query}" # 模拟浏览器请求头,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(api_url, headers=headers) data = response.json() # 提取首个产品URL if data.get("products"): first_product = data["products"][0] product_url = f"https://www.sobrico.com/p/{first_product['url_key']}.html" print(product_url) else: print("未找到搜索结果")
方法2:使用无头浏览器渲染JS
若API接口难以定位或存在反爬限制,可使用Selenium、Playwright等工具模拟浏览器执行JS,获取完整渲染后的页面内容。
示例代码(Selenium):
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options import time search_query = "2608664131" url = f"https://www.sobrico.com/#Prod_Live_Sobrico%5Bquery%5D={search_query}" # 配置无头模式浏览器 chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 等待页面渲染完成 time.sleep(2) # 提取首个产品链接 try: first_link = driver.find_element(By.CSS_SELECTOR, ".product-item a") print(first_link.get_attribute("href")) except: print("未找到产品链接") driver.quit()
注意事项
- 请求API时需携带合规的请求头,避免被网站识别为爬虫。
- 无头浏览器方法效率较低,适合小规模抓取,大规模使用易触发反爬机制。
内容的提问来源于stack exchange,提问作者Jean Besin
相关产品推荐
相关产品推荐

