如何用Python抓取网页数字?代码返回空变量求助
问题解决指南
尝试从目标网站抓取页面上的数字,但代码返回空变量,以下是问题分析和修复方案:
问题代码
import requests from bs4 import BeautifulSoup url = "https://www.quantalys.com/Recherche?Values.lstCategorie=27&Values.lstIdProduits=1&Values.bExcludeUncommercialized=true&Values.bETF=true" soup = BeautifulSoup(requests.get(url).content, "html.parser") nbre = soup.find("DataTables_Table_0_info") nbre
相关页面截图:
问题原因
- 选择器使用错误:
DataTables_Table_0_info是目标元素的id属性值,你直接把它作为标签名传给find()方法,自然找不到元素。find()方法第一个参数默认匹配HTML标签(如div、p),要通过id查找需指定对应参数。 - 可能的动态加载:该页面的表格数据可能由JavaScript动态渲染,
requests获取的是静态HTML源码,可能不包含目标元素。
修复方案
方案1:修正选择器(先验证静态HTML是否包含元素)
修改代码,正确通过id定位元素:
import requests from bs4 import BeautifulSoup url = "https://www.quantalys.com/Recherche?Values.lstCategorie=27&Values.lstIdProduits=1&Values.bExcludeUncommercialized=true&Values.bETF=true" response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") # 方法1:通过id参数查找 nbre = soup.find(id="DataTables_Table_0_info") # 方法2:使用CSS选择器(等价) # nbre = soup.select_one("#DataTables_Table_0_info") if nbre: # 提取文本中的数字,比如类似"显示1到10,共25条"这样的内容 print(nbre.text.strip()) else: print("未找到目标元素,大概率是数据动态加载导致")
方案2:处理动态加载内容
如果静态HTML里确实没有目标元素,就需要用浏览器自动化工具模拟页面加载:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url = "https://www.quantalys.com/Recherche?Values.lstCategorie=27&Values.lstIdProduits=1&Values.bExcludeUncommercialized=true&Values.bETF=true" # 初始化Chrome浏览器,需提前安装ChromeDriver并配置环境变量 driver = webdriver.Chrome() driver.get(url) # 等待目标元素加载完成,超时时间10秒 nbre_element = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "DataTables_Table_0_info")) ) # 提取文本并打印 print(nbre_element.text.strip()) driver.quit()
内容的提问来源于stack exchange,提问作者jacques
相关产品推荐
相关产品推荐

