You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用URL列表时Python-Requests脚本重复下载同一页面的问题

问题描述

作为某网站合法所有者,我需要下载站内图片,已将各页面URL存入urls.txt:

http://www.x10.com.cn/front/active-view-content?code=pRqKeo&active_role_id=p36Pao&content_id=p36Pao
http://www.x10.com.cn/front/active-view-content?code=pRqKeo&active_role_id=p36Pao&content_id=zZB9ez
http://www.x10.com.cn/front/active-view-content?code=pRqKeo&active_role_id=p36Pao&content_id=zAKnqp
http://www.x10.com.cn/front/active-view-content?code=pRqKeo&active_role_id=p36Pao&content_id=227420
http://www.x10.com.cn/front/active-view-content?code=pRqKeo&active_role_id=p36Pao&content_id=227378
--snip--

使用以下Python脚本下载时,始终重复下载第一个页面的内容,尽管打印输出显示URL每次都不同,添加time.sleep(1)也无效:

import requests
import bs4
import time

with open("urls.txt") as file_object:
    urls = file_object.readlines()
for url in urls:
    print(f"URL={url}")
    resp = requests.get(url)
    txt = resp.text
    sp = bs4.BeautifulSoup(txt)
    names = sp.select("div[class='field-value xingming']")
    imgs = sp.select("img[data-src]")
    for i in range(0, 2):
        img = imgs[i]
        link = img.get("data-src")
        res2 = requests.get(link)
        with open(f"caches/{names[0].text}_{i + 1}.jpg", "wb") as file_object:
            for chunk in res2.iter_content(100000):
                file_object.write(chunk)
            print(f"DOWNLOADING {names[0]}_{i + 1}.jpg")
    time.sleep(1)
问题排查与修复

核心原因:URL包含换行符

readlines()会保留每行末尾的换行符(\n),导致实际请求的URL是带换行的格式(比如"http://xxx.com\n")。服务器收到这种无效URL时,可能会重定向到默认页面或返回固定内容,最终导致每次都下载第一个页面的资源。

修复步骤

  1. 清理URL的空白字符:在循环中对每个URL执行strip(),去掉首尾的换行、空格等无效字符,同时跳过空行:

    for url in urls:
        url = url.strip()
        if not url:
            continue
        print(f"URL={url}")
        resp = requests.get(url)
    
  2. 指定BeautifulSoup解析器:未指定解析器可能导致HTML解析不一致,建议明确指定html.parser:

    sp = bs4.BeautifulSoup(txt, 'html.parser')
    
  3. 增加请求状态校验:确保请求成功后再处理内容,避免使用失败请求的返回结果:

    resp = requests.get(url)
    resp.raise_for_status()  # 请求失败时直接抛出异常
    txt = resp.text
    
  4. 避免索引越界风险:检查names和imgs是否有足够元素,避免后续代码报错:

    if not names or not imgs or len(imgs) < 2:
        print(f"页面{url}缺少必要元素,跳过")
        continue
    

修复后的完整代码

import requests
import bs4
import time
import os

# 确保caches文件夹存在,避免写入时出错
os.makedirs("caches", exist_ok=True)

with open("urls.txt") as file_object:
    urls = file_object.readlines()

for url in urls:
    url = url.strip()
    if not url:
        continue
    print(f"URL={url}")
    try:
        resp = requests.get(url)
        resp.raise_for_status()
        txt = resp.text
        sp = bs4.BeautifulSoup(txt, 'html.parser')
        names = sp.select("div[class='field-value xingming']")
        imgs = sp.select("img[data-src]")
        
        if not names or not imgs or len(imgs) < 2:
            print(f"页面{url}缺少必要元素,跳过")
            continue
            
        name = names[0].text.strip()
        for i in range(0, 2):
            img = imgs[i]
            link = img.get("data-src")
            if not link:
                print(f"第{i+1}张图片无有效链接,跳过")
                continue
            res2 = requests.get(link)
            res2.raise_for_status()
            filename = f"caches/{name}_{i + 1}.jpg"
            with open(filename, "wb") as file_object:
                for chunk in res2.iter_content(100000):
                    file_object.write(chunk)
            print(f"DOWNLOADING {filename}")
    except Exception as e:
        print(f"处理URL {url}时出错: {str(e)}")
    time.sleep(1)

内容的提问来源于stack exchange,提问作者Benny Frog

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 05:37:12