Python中批量写入论坛多页图片链接至TXT时的覆盖问题求助
问题原因
你代码的核心问题是每次调用list_image_links函数时,都用'w'(写入模式)打开文件——这个模式会清空文件原有内容后重新写入,所以新页面的链接会直接覆盖之前的内容。
解决方案
有两种高效的修改方式:
方式一:改用追加模式写入
先在循环外初始化文件、写入开头标识,之后每次爬取页面时用'a'(追加模式)写入新链接,避免覆盖原有内容。
修改后的代码:
from bs4 import BeautifulSoup import requests def list_image_links(url): response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") image_links = [] for link in soup.find_all('a'): href = link.get('href') if href is not None and 'attach' in href and not href.endswith('image'): image_links.append(href) # 追加模式写入新链接 with open('my_file.txt', 'a') as my_file: for branch in image_links: my_file.write(branch + '\n') print(f"已追加{len(image_links)}条链接到文件") # 初始化文件(仅执行一次) with open('my_file.txt', 'w') as my_file: my_file.write('image links:' + '\n') # 遍历所有页面 i = 0 while i <= 5175: list_image_links(f'https://forum.ubuntu.ir/index.php?topic=211.{i}') i += 15
方式二:先收集所有链接,再统一写入
这种方式减少频繁的文件IO操作,先把所有页面的链接收集到一个列表里,最后一次性写入文件,效率更高。
修改后的代码:
from bs4 import BeautifulSoup import requests def get_image_links(url): response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") image_links = [] for link in soup.find_all('a'): href = link.get('href') if href is not None and 'attach' in href and not href.endswith('image'): image_links.append(href) return image_links # 收集所有页面的链接 all_image_links = [] i = 0 while i <= 5175: page_links = get_image_links(f'https://forum.ubuntu.ir/index.php?topic=211.{i}') all_image_links.extend(page_links) print(f"已抓取第{i//15 +1}页,新增{len(page_links)}条链接") i += 15 # 统一写入文件 with open('my_file.txt', 'w') as my_file: my_file.write('image links:' + '\n') for link in all_image_links: my_file.write(link + '\n') print(f"所有链接已写入文件,共{len(all_image_links)}条")
额外提示
- 建议给请求添加浏览器模拟头,避免被论坛反爬拦截:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers)
- 可以添加异常处理逻辑,避免单个页面请求失败导致整个程序中断。
内容的提问来源于stack exchange,提问作者Mr Hunter
相关产品推荐
相关产品推荐

