You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

For循环调用BeautifulSoup爬虫函数仅返回最后一条内容,其余为空求排查

爬虫仅返回最后一个URL内容的问题排查

我编写了一个接收URL、速率限制(ratelimit)和目标元素选择器三个参数的html_scraper爬虫函数,通过for循环遍历url_list中的每个URL调用该函数,期望获取每个页面的爬取内容,但仅最后一个URL的内容能正确返回,前两个的pageTitle和content均为None。用于收集结果的scraped_html列表定义在for循环外部,请问问题出在哪里?

完整代码及输出

from bs4 import BeautifulSoup
import lxml
import requests

scraped_html = []

def html_scraper(url, ratelimit=1.0, target_element_selector='example-id-0'):

    page_contents = {
        'pageTitle': None,
        'content': None
    }

    result = requests.get(url).text

    soup = BeautifulSoup(result, 'lxml')

    page_contents['pageTitle'] = soup.find('h1')

    page_contents['content'] = soup.find(id=target_element_selector)
    
    time.sleep(ratelimit)

    return page_contents


if __name__ == '__main__':
    
    url_list = [
        'https://example.com/page-1',
        'https://example.com/page-2',
        'https://example.com/page-3',
    ]

    for url in url_list:
            try:
                scraped = html_scraper(url, 0.5, 'example-id-1')

                scraped_html.append(scraped)

            except Exception as e:
                print(e)
    
     print(scraped_html)
     # [{'pageTitle': None, 'content': None}, {'pageTitle': None, 'content': None}, {'pageTitle': Example Page 3 Title, 'content': <div id="example-id-1">Blah-blah-blah-blah-blah</div>}]

问题原因及修复方案

核心问题点

  1. 缺失time模块导入:函数中使用了time.sleep()但未导入time模块,虽然被try-except捕获,但会导致函数执行报错,影响程序稳定性。
  2. 未校验HTTP请求状态:requests.get(url)可能返回404、500等错误响应,此时直接取.text得到的是错误页面内容,自然无法找到目标元素,导致返回None。
  3. 页面结构不匹配:前两个页面的实际结构与第三个不一致:
    • 前两个页面没有<h1>标签,或者<h1>标签的存在形式不同(比如带特殊class/属性)
    • 前两个页面不存在ID为example-id-1的元素,或者元素ID与传入的选择器不匹配

修复步骤

  • 补上time模块导入:在代码开头添加import time
  • 校验HTTP响应状态,避免错误页面干扰:
    response = requests.get(url)
    response.raise_for_status()  # 遇到HTTP错误时抛出异常
    result = response.text
    
  • 优化元素查找逻辑,获取文本而非Tag对象(同时增强容错):
    # 获取h1文本,不存在则返回None
    page_contents['pageTitle'] = soup.find('h1').get_text(strip=True) if soup.find('h1') else None
    # 获取目标元素文本,不存在则返回None
    target_element = soup.find(id=target_element_selector)
    page_contents['content'] = target_element.get_text(strip=True) if target_element else None
    
  • 完善异常捕获,打印具体错误信息便于排查:
    except Exception as e:
        print(f"处理URL {url}时出错: {str(e)}")
    

内容的提问来源于stack exchange,提问作者japonix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 21:13:19