You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中批量写入论坛多页图片链接至TXT时的覆盖问题求助

问题原因

你代码的核心问题是每次调用list_image_links函数时,都用'w'(写入模式)打开文件——这个模式会清空文件原有内容后重新写入,所以新页面的链接会直接覆盖之前的内容。

解决方案

有两种高效的修改方式:

方式一:改用追加模式写入

先在循环外初始化文件、写入开头标识,之后每次爬取页面时用'a'(追加模式)写入新链接,避免覆盖原有内容。

修改后的代码:

from bs4 import BeautifulSoup
import requests

def list_image_links(url):
    response = requests.get(url)
    soup = BeautifulSoup(response.content, "html.parser")
    
    image_links = []
    for link in soup.find_all('a'):
        href = link.get('href')
        if href is not None and 'attach' in href and not href.endswith('image'):
            image_links.append(href)
    
    # 追加模式写入新链接
    with open('my_file.txt', 'a') as my_file:
        for branch in image_links:
            my_file.write(branch + '\n')
    print(f"已追加{len(image_links)}条链接到文件")

# 初始化文件(仅执行一次)
with open('my_file.txt', 'w') as my_file:
    my_file.write('image links:' + '\n')

# 遍历所有页面
i = 0
while i <= 5175:
    list_image_links(f'https://forum.ubuntu.ir/index.php?topic=211.{i}')
    i += 15

方式二:先收集所有链接,再统一写入

这种方式减少频繁的文件IO操作,先把所有页面的链接收集到一个列表里,最后一次性写入文件,效率更高。

修改后的代码:

from bs4 import BeautifulSoup
import requests

def get_image_links(url):
    response = requests.get(url)
    soup = BeautifulSoup(response.content, "html.parser")
    
    image_links = []
    for link in soup.find_all('a'):
        href = link.get('href')
        if href is not None and 'attach' in href and not href.endswith('image'):
            image_links.append(href)
    return image_links

# 收集所有页面的链接
all_image_links = []
i = 0
while i <= 5175:
    page_links = get_image_links(f'https://forum.ubuntu.ir/index.php?topic=211.{i}')
    all_image_links.extend(page_links)
    print(f"已抓取第{i//15 +1}页,新增{len(page_links)}条链接")
    i += 15

# 统一写入文件
with open('my_file.txt', 'w') as my_file:
    my_file.write('image links:' + '\n')
    for link in all_image_links:
        my_file.write(link + '\n')
print(f"所有链接已写入文件,共{len(all_image_links)}条")
额外提示
  • 建议给请求添加浏览器模拟头,避免被论坛反爬拦截:
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
response = requests.get(url, headers=headers)
  • 可以添加异常处理逻辑,避免单个页面请求失败导致整个程序中断。

内容的提问来源于stack exchange,提问作者Mr Hunter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 05:58:26