使用BS4+Requests爬取网页标题遇AttributeError问题求助
报错原因与解决方法
报错原因
- 短链接跳转未获取到最终页面:你使用的短链接存在多次跳转,若跳转包含JavaScript触发的页面跳转,
requests无法执行JS,只能获取到跳转中间页的HTML,中间页没有h1.entry-title标签,导致soup.find()返回None,调用.text时触发AttributeError。 - 未做元素存在性校验:代码直接调用
article_title.text,未先判断是否成功找到目标元素,一旦找不到就会报错。 - 目标页面结构可能变更:即使拿到最终页面,原代码依赖的
h1.entry-title选择器可能已失效,页面标题的标签或类名发生了变化。
解决方法
1. 确保获取到最终跳转页面
requests默认自动处理HTTP 30x重定向,但遇到JS跳转时需要手动处理。可以先打印请求后的实际URL,确认是否到达目标页面:
import requests import bs4 def main(url): headers = {"User-Agent": "Mozilla/5.0 (X11; Linux x86_64; rv:60.0) Gecko/20100101 Firefox/60.0"} response = requests.get(url, headers=headers, allow_redirects=True) print("实际请求URL:", response.url) # 查看是否是最终目标页面 data = response.text soup = bs4.BeautifulSoup(data, 'html.parser') # 后续代码...
如果打印的URL不是最终页面,说明是JS跳转,此时可以改用selenium模拟浏览器执行JS获取最终页面:
from selenium import webdriver from selenium.webdriver.firefox.options import Options import bs4 def main(url): options = Options() options.add_argument('--headless') # 无头模式,不弹出浏览器 driver = webdriver.Firefox(options=options) driver.get(url) driver.implicitly_wait(10) # 等待跳转完成 data = driver.page_source driver.quit() soup = bs4.BeautifulSoup(data, 'html.parser') # 后续代码...
2. 修正标题选择器并增加异常校验
拿到最终页面的HTML后,通过浏览器F12查看页面源代码,找到标题对应的标签和属性,修正选择器,同时增加判断避免报错:
# 替换为实际页面的标题选择器 article_title = soup.find('h1', class_='post-title') if article_title: print(article_title.get_text(strip=True)) # get_text更安全,strip去除多余空格 else: print("未找到标题元素")
3. 完整修正后的代码示例(基于requests处理HTTP重定向)
import requests import bs4 def main(url): headers = {"User-Agent": "Mozilla/5.0 (X11; Linux x86_64; rv:60.0) Gecko/20100101 Firefox/60.0"} try: response = requests.get(url, headers=headers, allow_redirects=True, timeout=10) response.raise_for_status() # 检查请求是否成功 print("实际访问URL:", response.url) soup = bs4.BeautifulSoup(response.text, 'html.parser') # 替换为实际页面的标题选择器 article_title = soup.find('h1', class_='entry-title') if article_title: print("文章标题:", article_title.get_text(strip=True)) else: print("未找到标题元素,请检查页面结构") # 获取内容 content = soup.find_all(attrs={'class':'td-post-content'}) if content: for part in content: print(part.get_text(strip=True)) else: print("未找到内容元素") except requests.exceptions.RequestException as e: print("请求出错:", e) if __name__=='__main__': url = "https://shorturl.at/fgLU8" main(url)
内容的提问来源于stack exchange,提问作者PRITAM BHAKTA
相关产品推荐
相关产品推荐

