You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量抓取Excel文件中所有URL并按URL_ID命名保存?

解决方案

步骤1:安装必要的Python库

打开终端运行以下命令:

pip install pandas requests beautifulsoup4 openpyxl

(openpyxl是读取.xlsx文件的依赖库)

步骤2:完整代码实现

import pandas as pd
import requests
from bs4 import BeautifulSoup
import os

# 创建保存文本文件的文件夹(不存在则自动创建)
output_folder = "scraped_content"
if not os.path.exists(output_folder):
    os.makedirs(output_folder)

# 读取Excel文件,假设文件内有两列:URL_ID(唯一标识)和URL(待抓取链接)
df = pd.read_excel("Input.xlsx", engine="openpyxl")

# 循环处理每一条链接
for index, row in df.iterrows():
    url_id = row["URL_ID"]
    target_url = row["URL"]
    
    try:
        # 发送GET请求获取网页内容,设置超时避免程序卡住
        response = requests.get(target_url, timeout=10)
        response.raise_for_status()  # 请求失败时直接抛出异常
        
        # 解析HTML提取纯文本内容
        soup = BeautifulSoup(response.text, "html.parser")
        # 提取页面纯文本(若目标网页有特定正文容器,可修改此处精准定位)
        page_content = soup.get_text(strip=True, separator="\n")
        
        # 保存到以URL_ID命名的文本文件
        save_path = os.path.join(output_folder, f"{url_id}.txt")
        with open(save_path, "w", encoding="utf-8") as f:
            f.write(page_content)
        
        print(f"完成:{url_id}")
    
    except Exception as e:
        # 捕获异常并打印,不终止整个程序
        print(f"{url_id}处理失败:{str(e)}")
        continue

关键说明

  • Excel格式要求:确保Input.xlsx包含URL_ID和URL两列,130条数据对应填充即可。
  • 内容提取优化:如果目标网页的正文在特定标签内(比如<div class="article-body">),可以把提取代码改成soup.find("div", class_="article-body").get_text(strip=True, separator="\n"),提升内容精准度。
  • 反爬注意:频繁请求可能触发网站反爬机制,可添加请求头模拟浏览器:
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"}
    response = requests.get(target_url, headers=headers, timeout=10)
    
  • 编码问题:保存文件用utf-8编码,避免中文乱码。

内容的提问来源于stack exchange,提问作者kingslayer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 03:10:22