You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从CSV文件读取URL列表用于Python网页爬取迭代

从CSV批量读取URL执行网页爬取的解决方案

实现步骤

不需要用字典,直接读取CSV每行的URL并遍历执行你的爬取逻辑即可,以下是具体实现:

1. 准备URL列表CSV

确保你的CSV文件每行仅包含一个目标URL,示例urls.csv内容:

https://example.com/page1
https://example.com/page2
https://example.com/page3

2. 修改爬取代码适配批量URL

假设你原有固定URL的爬取代码类似下面的结构,我们直接改造为批量读取版本:

原有固定URL代码(示例)

import requests
from bs4 import BeautifulSoup

fixed_url = "https://example.com"
response = requests.get(fixed_url)
soup = BeautifulSoup(response.text, "html.parser")
# 你的爬取逻辑,比如提取页面标题
title = soup.title.string
print(title)

改造后的批量爬取代码

import requests
from bs4 import BeautifulSoup
import csv
import time

# 读取CSV中的URL并遍历处理
with open("urls.csv", "r", encoding="utf-8") as csv_file:
    # 逐行读取CSV(每行一个URL)
    url_list = csv.reader(csv_file)
    for row in url_list:
        url = row[0].strip()
        # 跳过空行
        if not url:
            continue
        
        try:
            # 执行爬取逻辑,加入超时避免请求挂起
            response = requests.get(url, timeout=10)
            # 触发HTTP错误(比如404、500)
            response.raise_for_status()
            
            soup = BeautifulSoup(response.text, "html.parser")
            # 替换为你实际的爬取操作(提取数据、写入文件等)
            page_title = soup.title.string if soup.title else "无页面标题"
            print(f"处理完成: {url} | 标题: {page_title}")
            
            # 可选:添加请求延迟,避免触发目标网站反爬机制
            time.sleep(1)
            
        except Exception as e:
            # 捕获异常,单个URL失败不影响整体任务
            print(f"处理失败 {url}: {str(e)}")

关键注意点

  • 异常捕获:必须加入,防止单个URL的请求错误导致整个程序终止
  • 超时设置:避免请求长时间卡住占用资源
  • 请求延迟:URL数量较多时建议添加,降低被目标网站封禁的风险
  • 空行处理:跳过CSV中的空行,避免无效请求

内容的提问来源于stack exchange,提问作者Orochimaru Sama

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 20:25:23