如何使用bs4抓取多个URL内容并分别保存到独立文本文件
问题原因
你的问题核心是代码缩进错误:当前写入文件的with代码块完全位于for循环外部,只会等所有URL遍历完成后执行一次。此时循环里的name和text变量存储的都是最后一次循环的取值,自然只会生成最后一个页面的保存文件。
修复方案
场景1:每个页面单独保存为1个TXT文件
只需要把写入文件的代码缩进,放到for循环内部即可,每处理完一个页面就执行一次写入:
import bs4 import requests from fake_useragent import UserAgent import io import time # 可选,加请求延迟避免被站点封禁 urls = ['https://en.m.wikipedia.org/wiki/Grove_(nature)','https://en.wikipedia.org/wiki/Azadirachta_indica','https://en.wikipedia.org/wiki/Olive'] user_agent = UserAgent() for url in urls: page = requests.get(url, headers={"user-agent": user_agent.chrome}) tree = bs4.BeautifulSoup(page.text, 'html.parser') title = tree.find('title').get_text() text = tree.find_all('p')[1].get_text() name = title + '.txt' # 缩进后放入循环内部,每次处理完页面直接写入 with io.open(name, "w", encoding="utf-8") as text_file: text_file.write(text) print(f'{name} 已保存') time.sleep(1) # 可选,每次请求间隔1秒
场景2:所有页面内容保存到同一个TXT文件
把文件打开操作放在循环外部,循环内部仅执行写入动作即可:
import bs4 import requests from fake_useragent import UserAgent import io import time urls = ['https://en.m.wikipedia.org/wiki/Grove_(nature)','https://en.wikipedia.org/wiki/Azadirachta_indica','https://en.wikipedia.org/wiki/Olive'] user_agent = UserAgent() # 提前打开统一的存储文件 with io.open("所有页面内容.txt", "w", encoding="utf-8") as text_file: for url in urls: page = requests.get(url, headers={"user-agent": user_agent.chrome}) tree = bs4.BeautifulSoup(page.text, 'html.parser') title = tree.find('title').get_text() text = tree.find_all('p')[1].get_text() # 写入当前页面内容,加标识和换行分隔不同页面 text_file.write(f"===== {title} =====\n{text}\n\n") print(f'{title} 内容已写入') time.sleep(1) print('全部内容处理完成')
补充提醒
部分特殊页面可能不存在第二个p标签,直接取find_all('p')[1]会触发索引报错,可以先判断p标签列表长度再取值,代码健壮性会更高。
内容的提问来源于stack exchange,提问作者user16748482
相关产品推荐
相关产品推荐

