You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何用BeautifulSoup爬取网站百分比数值始终返回0?求解决方法

解决动态加载页面的百分比提取问题

你遇到的问题是因为目标网站的百分比数据是通过JavaScript动态渲染的,requests库只能获取页面的静态初始HTML,此时这些百分比还没被JS替换成真实数值,所以拿到的是默认的0%。

以下是两种可行的解决方案:

方法一:使用Selenium模拟浏览器加载(快速实现)

Selenium可以模拟真实浏览器打开页面,等待JavaScript执行完成后再提取数据。

  1. 先安装依赖:
pip install selenium
  1. 下载对应浏览器的驱动(比如Chrome的ChromeDriver,需与浏览器版本匹配),确保驱动路径能被Python识别。

代码示例:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

URL = "https://www.horoscope.fr/horoscopes/aujourdhui/scorpion"

# 初始化Chrome浏览器驱动
driver = webdriver.Chrome()
driver.get(URL)

# 等待页面加载完成,直到strong标签出现含%的内容(最多等10秒)
try:
    WebDriverWait(driver, 10).until(
        EC.text_to_be_present_in_element((By.TAG_NAME, "strong"), "%")
    )
    # 提取所有带百分比的strong标签内容
    trucs = driver.find_elements(By.TAG_NAME, "strong")
    for truc in trucs:
        text = truc.text
        if "%" in text:
            print(text)
finally:
    # 关闭浏览器
    driver.quit()

方法二:直接调用API接口(更高效)

打开浏览器开发者工具(F12),切换到「网络」标签,刷新页面后筛选XHR请求,能找到返回每日运势数据的API接口。直接请求该接口即可拿到JSON格式的原始数据,无需处理页面渲染。

代码示例(需根据实际API地址调整):

import requests

# 替换为你在开发者工具中找到的真实API地址
API_URL = "https://www.horoscope.fr/api/horoscope/daily/scorpion"
response = requests.get(API_URL)
data = response.json()

# 根据API返回的JSON结构提取百分比(示例结构,需自行验证)
if "scores" in data:
    for category, score in data["scores"].items():
        print(f"{category}: {score}%")

内容的提问来源于stack exchange,提问作者sukiheaven

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 23:10:24