You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取返回空HTML标签:无法获取目标网站艺术品数据求助

解决爬取centerforbookarts.org/book-shop数据为空的问题

问题原因分析

你遇到的情况大概率是以下两种原因:

  • 内容动态渲染:网站通过JavaScript加载艺术品数据,requests获取的是初始HTML源码,还没执行JS渲染,所以找不到目标标签;
  • 反爬拦截:网站检测到请求不是来自真实浏览器,返回的内容不完整;
  • 若Selenium也失效,可能是没有等待页面完全加载就获取源码。

针对性解决方案

1. 给requests添加请求头模拟浏览器

很多网站会校验User-Agent字段,添加后可绕过基础反爬:

from bs4 import BeautifulSoup
import requests

url = "https://centerforbookarts.org/book-shop"
# 模拟Chrome浏览器的请求头,可根据自己浏览器版本修改
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

response = requests.get(url, headers=headers)
# 先打印响应文本,确认是否包含目标内容
# print(response.text)
soup = BeautifulSoup(response.text, "lxml")
element = soup.find_all("section", {"class": "posts"})
print(element)

2. 用Selenium等待页面完全加载

如果是动态渲染问题,Selenium需要等待JS执行完毕、目标元素出现后再获取源码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

url = "https://centerforbookarts.org/book-shop"

# 初始化浏览器驱动(确保ChromeDriver已配置到环境变量,或指定路径)
driver = webdriver.Chrome()
driver.get(url)

try:
    # 最多等待10秒,直到目标section元素出现
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "posts"))
    )
    # 获取渲染后的完整页面源码
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, "lxml")
    element = soup.find_all("section", {"class": "posts"})
    print(element)
finally:
    # 不管成功失败都关闭浏览器
    driver.quit()

3. 极端情况:应对Cloudflare等强反爬

如果上述方法都失效,可能网站用了Cloudflare之类的反爬验证,此时可以使用undetected-chromedriver库(它能绕过大部分浏览器指纹检测):

from undetected_chromedriver import Chrome
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

url = "https://centerforbookarts.org/book-shop"

driver = Chrome()
driver.get(url)

try:
    WebDriverWait(driver, 15).until(
        EC.presence_of_element_located((By.CLASS_NAME, "posts"))
    )
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, "lxml")
    element = soup.find_all("section", {"class": "posts"})
    print(element)
finally:
    driver.quit()

使用前需要先安装库:pip install undetected-chromedriver

内容的提问来源于stack exchange,提问作者toru

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 21:08:31