使用Python BeautifulSoup爬取网页返回空列表问题求助
问题分析与解决方法
可能原因
- 反爬拦截:目标网站检测到请求来自Python爬虫(默认
requests的User-Agent为python-requests/x.x.x),返回的HTML内容与浏览器实际加载的不一致,导致无法找到目标标签。 - 动态内容加载:页面中
class="t3t1"的td标签数据是通过JavaScript动态渲染生成的,requests只能获取静态HTML源码,无法执行JS,因此拿不到动态加载的数据。
解决步骤
步骤1:添加请求头模拟浏览器请求
先尝试给requests.get添加浏览器的User-Agent,绕过基础反爬:
import requests from bs4 import BeautifulSoup headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } reqget = requests.get('https://fubon-ebrokerdj.fbs.com.tw/z/zg/zg_A_0_5.djhtm', headers=headers) # 检查请求是否成功 print(reqget.status_code) # 验证返回的HTML中是否包含目标class print('t3t1' in reqget.text) html = reqget.text sp = BeautifulSoup(html, 'html.parser') data = [i.text.strip() for i in sp.find_all('td', class_='t3t1')] print(data)
如果print('t3t1' in reqget.text)返回True,说明请求到了正确内容,此时应该能输出目标数据;如果返回False,则说明是动态加载的问题。
步骤2:使用浏览器自动化工具获取动态内容
若数据是JS动态渲染的,需要用selenium模拟浏览器加载页面(需提前安装selenium和对应浏览器驱动,比如ChromeDriver):
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import time # 配置无头模式(可选,不弹出浏览器窗口) chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--disable-gpu') driver = webdriver.Chrome(options=chrome_options) driver.get('https://fubon-ebrokerdj.fbs.com.tw/z/zg/zg_A_0_5.djhtm') # 等待页面加载完成(可根据实际情况调整等待时间) time.sleep(3) html = driver.page_source sp = BeautifulSoup(html, 'html.parser') data = [i.text.strip() for i in sp.find_all('td', class_='t3t1')] print(data) driver.quit()
额外提示
- 可以先打印
reqget.text查看返回的HTML内容,确认是否和浏览器开发者工具中看到的一致,这是排查问题的关键。 - 如果网站有更严格的反爬机制(如验证码、Cookie验证),可能需要进一步处理Cookie或使用代理,但优先从上述两种情况入手排查。
内容的提问来源于stack exchange,提问作者0983
相关产品推荐
相关产品推荐

