You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup从URL列表爬取合并数据仅返回首个结果及报错如何解决

问题原因和解决方法

仅能获取单州数据的原因

你的字段提取代码写在了URL请求循环的外部,请求循环执行时会每次覆盖doc变量的取值,循环结束后doc仅保留最后一次请求返回的解析结果。如果你实际拿到的是首个州123的结果,大概率是后两个URL请求触发了反爬、返回了无效页面,doc没有被后续请求的结果覆盖,才会只能拿到单州数据。

AttributeError报错的原因

你调整代码后把所有解析后的BeautifulSoup对象存入了doc列表,而select是BeautifulSoup实例的专属方法,列表本身没有该方法,直接调用doc.select自然会报错,确实需要新增一层循环遍历doc列表里的每个解析对象再处理。

两种修正方案

方案1:边请求边处理(更节省内存,推荐)

直接把字段提取逻辑放到请求循环内部,不需要额外存储所有页面的解析结果:

states = ['123', '124', '125']
urls = []
for state in states:
    url = f'www.something.com/geo={state}'
    urls.append(url)

rows = []
# 客户端实例化放到循环外,不需要每次请求都新建实例,节省资源
client = ScrapingBeeClient(api_key="API_KEY")
for url in urls:
    response = client.get(url)
    # 新增响应状态校验,避免处理错误返回内容
    if response.status_code != 200:
        print(f"请求{url}失败,状态码:{response.status_code}")
        continue
    doc = BeautifulSoup(response.text, 'html.parser')
    listings = doc.select('.is-9-desktop')
    for listing in listings:
        row = {}
        try:
            row['name'] = listing.select_one('.result-title').text.strip()
        except:
            print("no name")
        try:
            row['add'] = listing.select_one('.address-text').text.strip()
        except:
            print("no add")
        try:
            row['mention'] = listing.select_one('.review-mention-block').text.strip()
        except:
            pass
        rows.append(row)

方案2:先存所有解析结果再批量处理

如果你确实需要先把所有页面解析结果存储下来再统一处理,可以新增一层循环遍历解析结果列表:

states = ['123', '124', '125']
urls = []
for state in states:
    url = f'www.something.com/geo={state}'
    urls.append(url)

docs = []
client = ScrapingBeeClient(api_key="API_KEY")
for url in urls:
    response = client.get(url)
    if response.status_code == 200:
        docs.append(BeautifulSoup(response.text, 'html.parser'))

rows = []
# 新增一层循环遍历每个页面的解析结果
for doc in docs:
    listings = doc.select('.is-9-desktop')
    for listing in listings:
        row = {}
        try:
            row['name'] = listing.select_one('.result-title').text.strip()
        except:
            print("no name")
        try:
            row['add'] = listing.select_one('.address-text').text.strip()
        except:
            print("no add")
        try:
            row['mention'] = listing.select_one('.review-mention-block').text.strip()
        except:
            pass
        rows.append(row)

内容的提问来源于stack exchange,提问作者user17593319

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 14:24:01