You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BS4+Requests爬取网页标题遇AttributeError问题求助

报错原因与解决方法

报错原因

  • 短链接跳转未获取到最终页面:你使用的短链接存在多次跳转,若跳转包含JavaScript触发的页面跳转,requests无法执行JS,只能获取到跳转中间页的HTML,中间页没有h1.entry-title标签,导致soup.find()返回None,调用.text时触发AttributeError。
  • 未做元素存在性校验:代码直接调用article_title.text,未先判断是否成功找到目标元素,一旦找不到就会报错。
  • 目标页面结构可能变更:即使拿到最终页面,原代码依赖的h1.entry-title选择器可能已失效,页面标题的标签或类名发生了变化。

解决方法

1. 确保获取到最终跳转页面

requests默认自动处理HTTP 30x重定向,但遇到JS跳转时需要手动处理。可以先打印请求后的实际URL,确认是否到达目标页面:

import requests
import bs4

def main(url):
    headers = {"User-Agent": "Mozilla/5.0 (X11; Linux x86_64; rv:60.0) Gecko/20100101 Firefox/60.0"}
    response = requests.get(url, headers=headers, allow_redirects=True)
    print("实际请求URL:", response.url)  # 查看是否是最终目标页面
    data = response.text
    soup = bs4.BeautifulSoup(data, 'html.parser')
    # 后续代码...

如果打印的URL不是最终页面,说明是JS跳转,此时可以改用selenium模拟浏览器执行JS获取最终页面:

from selenium import webdriver
from selenium.webdriver.firefox.options import Options
import bs4

def main(url):
    options = Options()
    options.add_argument('--headless')  # 无头模式,不弹出浏览器
    driver = webdriver.Firefox(options=options)
    driver.get(url)
    driver.implicitly_wait(10)  # 等待跳转完成
    data = driver.page_source
    driver.quit()
    soup = bs4.BeautifulSoup(data, 'html.parser')
    # 后续代码...

2. 修正标题选择器并增加异常校验

拿到最终页面的HTML后,通过浏览器F12查看页面源代码,找到标题对应的标签和属性,修正选择器,同时增加判断避免报错:

# 替换为实际页面的标题选择器
article_title = soup.find('h1', class_='post-title')
if article_title:
    print(article_title.get_text(strip=True))  # get_text更安全,strip去除多余空格
else:
    print("未找到标题元素")

3. 完整修正后的代码示例(基于requests处理HTTP重定向)

import requests
import bs4

def main(url):
    headers = {"User-Agent": "Mozilla/5.0 (X11; Linux x86_64; rv:60.0) Gecko/20100101 Firefox/60.0"}
    try:
        response = requests.get(url, headers=headers, allow_redirects=True, timeout=10)
        response.raise_for_status()  # 检查请求是否成功
        print("实际访问URL:", response.url)
        soup = bs4.BeautifulSoup(response.text, 'html.parser')
        
        # 替换为实际页面的标题选择器
        article_title = soup.find('h1', class_='entry-title')
        if article_title:
            print("文章标题:", article_title.get_text(strip=True))
        else:
            print("未找到标题元素,请检查页面结构")
        
        # 获取内容
        content = soup.find_all(attrs={'class':'td-post-content'})
        if content:
            for part in content:
                print(part.get_text(strip=True))
        else:
            print("未找到内容元素")
    except requests.exceptions.RequestException as e:
        print("请求出错:", e)

if __name__=='__main__':
    url = "https://shorturl.at/fgLU8"
    main(url)

内容的提问来源于stack exchange,提问作者PRITAM BHAKTA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 14:35:19