You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何BeautifulSoup无法返回目标网页的数据元素?

问题分析与解决思路

1. 先修复代码基础错误

你的代码存在两个明显的运行错误:

  • 未导入requests库,直接运行会抛出导入异常
  • 变量img_url未定义就直接调用打印,同样会报错

先补上基础逻辑:

import requests  # 补充导入requests库
from bs4 import BeautifulSoup

url = "https://www.hebban.nl/rank"
response = requests.get(url)

soup = BeautifulSoup(response.content, 'html.parser')

books = soup.find_all('div', class_='row-fluid')
for book in books:
    title = book.find('a', class_='neutral').text.strip()
    author = book.find('span', class_='author').text.strip()
    img_url = book.find('img')['src']  # 补充图片URL提取逻辑

    print(title + ' by ' + author)
    print('Image URL: ' + img_url)

2. 大概率遭遇反爬拦截

直接用requests.get()发送的请求缺少浏览器标识,很容易被服务器拦截,返回空内容或错误状态码。你可以先打印response.status_code和response.text[:500],确认是否拿到了有效页面内容。

解决方法是添加模拟浏览器的请求头:

import requests
from bs4 import BeautifulSoup

url = "https://www.hebban.nl/rank"
# 添加浏览器UA标识,避免被识别为爬虫
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
response = requests.get(url, headers=headers)

# 先验证响应状态和内容
print(response.status_code)
print(response.text[:500])  # 打印前500字符,确认是否拿到页面

soup = BeautifulSoup(response.content, 'html.parser')

books = soup.find_all('div', class_='row-fluid')
for book in books:
    title = book.find('a', class_='neutral').text.strip()
    author = book.find('span', class_='author').text.strip()
    img_url = book.find('img')['src']

    print(title + ' by ' + author)
    print('Image URL: ' + img_url)

3. 检查页面DOM结构是否更新

如果加了请求头还是拿不到数据,可能是网站更新了页面元素的类名或结构。你可以手动打开页面,右键检查元素,确认书籍容器、标题、作者的选择器是否还是row-fluid、neutral、author——如果网站改版,这些选择器大概率会变化,需要重新调整。

4. 处理动态渲染的情况

如果页面内容是通过JavaScript异步加载的(比如滚动加载数据),直接用requests只能拿到静态HTML,无法获取渲染后的内容。这时候需要用工具模拟浏览器渲染:

以selenium为例:

from selenium import webdriver
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup
import time

# 需要提前安装ChromeDriver并配置环境变量
driver = webdriver.Chrome()
driver.get("https://www.hebban.nl/rank")
time.sleep(2)  # 等待页面加载完成

soup = BeautifulSoup(driver.page_source, 'html.parser')
# 这里的选择器需要根据实际页面DOM调整
books = soup.find_all('div', class_='row-fluid')

for book in books:
    title = book.find('a', class_='neutral').text.strip()
    author = book.find('span', class_='author').text.strip()
    img_url = book.find('img')['src']

    print(title + ' by ' + author)
    print('Image URL: ' + img_url)

driver.quit()

内容的提问来源于stack exchange,提问作者jsb92

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 16:55:23