You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取blocket.se仅返回head内容无body,求技术解决方案

问题原因及解决办法

核心原因

目标网站blocket.se存在反爬机制,你的requests请求被识别为非浏览器的爬虫请求,因此仅返回了不完整的页面内容(仅head部分)。而realpython.github.io/fake-jobs这类网站反爬规则宽松,所以你的代码能正常工作。

解决步骤

1. 完善请求头,模拟浏览器请求

默认requests请求缺少User-Agent字段,极易被识别为爬虫。添加浏览器的User-Agent可以伪装成正常访问:

import requests
from bs4 import BeautifulSoup

url = "https://www.blocket.se/annonser/hela_sverige?q=dunjacka"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# 尝试提取body内容
print(soup.body)

2. 若仍无效,说明页面是JS动态渲染的

部分网站内容通过JavaScript加载,requests只能获取初始HTML(可能仅含head),需用工具模拟浏览器渲染页面,比如selenium:

先安装依赖:

pip install selenium

编写代码:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup

url = "https://www.blocket.se/annonser/hela_sverige?q=dunjacka"

# 配置Chrome无头模式(不弹出浏览器窗口)
options = Options()
options.add_argument("--headless=new")
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=options)
driver.get(url)

# 获取渲染后的页面源码
page_source = driver.page_source
soup = BeautifulSoup(page_source, "html.parser")

# 提取body内容
print(soup.body)

driver.quit()

额外注意事项

  • 不要频繁发起请求,避免被封禁IP,可添加请求间隔(如time.sleep(2))。
  • 部分网站可能还会检测Cookie、Referer等字段,必要时可在headers中补充这些内容。

内容的提问来源于stack exchange,提问作者D Danne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 23:55:18