You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取商品价格返回0求助:如何提取目标网站的金价?

问题解决:正确提取目标网站商品价格

核心原因

你的请求未做浏览器伪装,网站识别为非人类访问,返回的HTML中价格被替换为0或未包含真实价格数据。此外原代码的insert用法冗余,且缺乏元素不存在的容错处理。

修正方案

方案1:添加请求头模拟浏览器访问

通过设置User-Agent等请求头,让网站认为是正常浏览器访问:

import requests
from bs4 import BeautifulSoup
from time import sleep


def maple_leaf_1(url, coins, n):
    # 模拟Chrome浏览器的请求头
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
        "Accept-Language": "en-US,en;q=0.9"
    }
    r = requests.get(url, headers=headers)
    soup = BeautifulSoup(r.text, "html.parser")
    sleep(3)
    
    # 先检查元素是否存在,避免报错
    coin_name_elem = soup.find("span", class_="base")
    coin_price_elem = soup.find("span", class_="price")
    
    if coin_name_elem and coin_price_elem:
        coin_name = coin_name_elem.text.strip()
        coin_price = coin_price_elem.text.strip()
        # 用extend简化列表添加操作
        coins[n].extend([url, coin_name, coin_price])
    else:
        coins[n].extend([url, "名称获取失败", "价格获取失败"])


def main():
    coins = [[]]
    maple_leaf_1("https://goldstocklive.com/1-oz-gold-2021-maple-leaf-coin.html", coins, 0)
    print(coins)

if __name__ == "__main__":
    main()

方案2:若价格为JS动态加载,使用Selenium

如果网站价格是通过JavaScript异步渲染的,requests无法获取真实数据,需用Selenium模拟完整浏览器加载:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from time import sleep


def maple_leaf_1(url, coins, n):
    # 配置无头Chrome模式,不弹出浏览器窗口
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")
    chrome_options.add_argument("--disable-gpu")
    chrome_options.add_argument("User-Agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    
    driver = webdriver.Chrome(options=chrome_options)
    driver.get(url)
    sleep(3)  # 等待页面及JS加载完成
    
    try:
        coin_name = driver.find_element(By.CLASS_NAME, "base").text.strip()
        coin_price = driver.find_element(By.CLASS_NAME, "price").text.strip()
        coins[n].extend([url, coin_name, coin_price])
    except Exception as e:
        coins[n].extend([url, "名称获取失败", f"价格获取失败: {str(e)}"])
    finally:
        driver.quit()


def main():
    coins = [[]]
    maple_leaf_1("https://goldstocklive.com/1-oz-gold-2021-maple-leaf-coin.html", coins, 0)
    print(coins)

if __name__ == "__main__":
    main()

注意事项

  • 控制请求频率,避免短时间内多次请求导致IP被封禁
  • 定期检查网站HTML结构,若元素的class或位置发生变化,需同步更新定位代码

内容的提问来源于stack exchange,提问作者bknapp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 13:25:38