You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取亚马逊问题时触发AttributeError,求解决方法

问题描述

我使用以下代码爬取亚马逊上的商品问题:

url = "https://www.amazon.com/ask/questions/asin/B0000CFLYJ/1/ref=ask_ql_psf_ql_hza?isAnswered=true"

r = requests.get("http://localhost:8050/render.html", params = {'url': url, 'wait': 3})
soup = BeautifulSoup(r.text, 'html.parser')
questions = soup.find_all('div', {'class':'a-fixed-left-grid-col a-col-right'})
print(questions)

question_list = []
for item in questions:
    question = item.find('a',{'class':'a-link-normal'}).text.strip()
    question_list.append(question)

但持续遇到如下错误:

AttributeError: 'NoneType' object has no attribute 'text'

我是否需要添加异常处理器?或者应该改用其他元素提取问题文本?我尝试过使用其下方的span元素,但没有效果:

<div class="a-fixed-left-grid-col a-col-right" style="padding-left:0%;float:left;">
<a class="a-link-normal" href="/ask/questions/Tx150GKDGF6FGAY/ref=ask_ql_ql_al_hza">
<span class="a-declarative" data-action="ask-no-op" data-ask-no-op='{"metricName":"top-question-text-click"}' data-csa-c-func-deps="aui-da-ask-no-op" data-csa-c-id="bsypsr-tzr1os-ttv9h7-td9hn6" data-csa-c-type="widget">
                
                  
                  
                    It comes in already made in a spray bottle, but yet says it's concentrated and gives dilution instructions? 

So do I use it as is, or dilute?
                  
                
              </span>

我目前仅尝试爬取第一页的问题,之后会循环处理其他页面,恳请各位提供帮助!


解决方案(基于@Unmitigated的回复修订,支持多页爬取)

感谢帮助!以下是修订后的代码,可实现多页商品问题爬取并导出至Excel:

question_list = []

# 定义爬取相关函数
def get_soup(url):
    # 调用Splash渲染页面
    r = requests.get("http://localhost:8050/render.html", params = {'url': url, 'wait': 3})
    soup = BeautifulSoup(r.text, 'html.parser')
    return soup

def get_questions(soup):
    for item in soup.select('.askTeaserQuestions > div'):
        question = item.find('a', {'class':'a-link-normal'}).getText(strip=True)
        question_list.append(question)

# 循环爬取页面(每页默认10条问题)
for x in range(1,6):
    soup = get_soup(f'https://www.amazon.com/ask/questions/asin/B0000CFLYJ/{x}')
    get_questions(soup)
    print(len(question_list))
    
    # 检测是否到达最后一页,是则终止循环
    if not soup.find('li',{'class':'a-disabled a-last'}):
        pass
    else:
        break

        
# 导出爬取结果至Excel文件
df = pd.DataFrame(question_list)
df.to_excel('SimpleGreen_Amazon_Questions_22oz_1pk_diff_seller.xlsx', index = False)
print('商品问题爬取及导出操作完成!')

内容的提问来源于stack exchange,提问作者mexicanRmy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 05:45:06