使用BeautifulSoup抓取atlasobscura.com时触发str无text属性报错如何解决
问题排查与修复方案
报错原因
报错触发在以下代码行:
href = 'https://www.atlasobscura.com' + str(s.find('div').find('a')['href'].text)
BeautifulSoup语法规则中,通过标签对象['属性名']获取的结果本身就是字符串类型的属性值,不属于Tag对象,没有.text属性,额外调用.text就会触发AttributeError: 'str' object has no attribute 'text'报错。
修复方案
1. 直接修改报错行
删除多余的.text调用即可,修改后代码如下:
href = 'https://www.atlasobscura.com' + s.find('div').find('a')['href']
注:['href']返回结果本身就是字符串,不需要额外用str()转换
2. 额外优化建议
- 将Excel导出逻辑挪到分页循环外,避免每次翻页都覆盖写入文件,减少IO消耗提升效率
- 增加异常捕获逻辑,避免单篇文章元素缺失导致整个爬虫终止
- 适当增加请求间隔,避免触发站点反爬限制
修复后的完整代码
import requests from bs4 import BeautifulSoup import pandas as pd import time title = [] link = [] date = [] headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} for x in range(1, 662): print(f'=====> 正在爬取第 {x} 页') url = f'https://www.atlasobscura.com/articles?page={x}' r = requests.get(url, headers=headers) r.encoding = r.apparent_encoding soup = BeautifulSoup(r.text, 'lxml') articles = soup.find_all('div', class_='col-md-4 col-sm-6 col-xs-12') for s in articles: try: story = s.find('div', class_='content-card-text').find('h3').find('span').text title.append(story) href = 'https://www.atlasobscura.com' + s.find('div').find('a')['href'] link.append(href) m_d_y = s.find('div', class_='detail-sm article-card-detail article-card-date').text.strip() date.append(m_d_y) print(story, href, m_d_y) except Exception as e: print(f"当前条目解析失败:{e}") continue # 每页爬完休息1秒,规避反爬 time.sleep(1) # 全部数据爬取完成后统一导出 atlasobscura = pd.DataFrame({ 'Title': title, 'Link': link, 'Date': date }) atlasobscura.to_excel('Atlasobscura.com.xlsx', index=False)
内容的提问来源于stack exchange,提问作者Peters7
相关产品推荐
相关产品推荐

