You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Requests爬取pixelford博客遇400错误及空列表问题

解决Pixelford博客爬虫400错误及标题提取问题

问题原因

  • 访问https://pixelford.com/blog/返回400错误,是因为服务器检测到请求并非来自标准浏览器——requests.get默认的请求头缺少User-Agent等标识,被服务器拦截。
  • 标题提取为空是因为请求失败后,返回的不是正常页面HTML,自然找不到目标标签;另外原代码中使用的class名称entry_title_link与实际页面不符。

修复方案

1. 添加请求头模拟浏览器

给requests.get添加headers参数,携带浏览器的User-Agent信息,让服务器认可请求来源。

2. 修正标签选择器

实际页面中文章标题的a标签class为entry-title-link(连字符连接),而非原代码中的下划线版本。

3. 增加请求状态校验

先判断请求状态码是否为200,确保请求成功后再进行HTML解析。

修改后的完整代码

import requests
from bs4 import BeautifulSoup

url = "https://pixelford.com/blog/"
# 模拟Chrome浏览器的请求头,可根据自己浏览器实际UA替换
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

response = requests.get(url, headers=headers)
if response.status_code == 200:
    soup = BeautifulSoup(response.text, 'html.parser')
    # 匹配正确的class名称
    title_links = soup.find_all('a', class_="entry-title-link")
    # 提取标题并去除多余空格
    titles = [link.get_text(strip=True) for link in title_links]
    print(titles)
else:
    print(f"请求失败,状态码:{response.status_code}")

额外提示

  • User-Agent可以在浏览器开发者工具的Network面板中查看:打开目标页面后,刷新并查看任意请求的Request Headers中的User-Agent字段。
  • 如果后续仍出现访问问题,可以尝试添加更多请求头字段(如Accept、Referer),进一步模拟浏览器行为。

内容的提问来源于stack exchange,提问作者David and Bethany

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 13:47:22