如何用Python将所有网页爬取数据写入单个HTML文件?
问题解决:循环数据无法全部保存到单个HTML文件
你的代码每次循环都用"w"模式打开output.html,这个模式会清空文件原有内容并写入新数据,所以最终只有最后一次循环的内容被保留。以下是两种可行的解决方案:
方案1:使用追加模式写入
将文件打开模式从"w"改为"a",同时建议在循环开始前先清空文件(避免多次运行代码时内容重复叠加):
import requests from bs4 import BeautifulSoup url = "https://gk-hindi.in/gk-questions?page=" # 先清空目标文件 with open("output.html", "w", encoding='utf-8') as file: pass i = 1 while i <= 48: req = requests.get(url + str(i)) soup = BeautifulSoup(req.content, "html.parser") mydivs = soup.find("div", {"class": "question-wrapper"}) print(mydivs) # 用追加模式写入内容 with open("output.html", "a", encoding='utf-8') as file: file.write(str(mydivs)) i += 1
方案2:先收集所有内容再一次性写入
这种方式减少了频繁的文件IO操作,效率更高:
import requests from bs4 import BeautifulSoup url = "https://gk-hindi.in/gk-questions?page=" all_content = "" i = 1 while i <= 48: req = requests.get(url + str(i)) soup = BeautifulSoup(req.content, "html.parser") mydivs = soup.find("div", {"class": "question-wrapper"}) print(mydivs) # 把每页内容追加到总字符串中 if mydivs: # 避免写入空内容 all_content += str(mydivs) i += 1 # 一次性写入所有收集到的内容 with open("output.html", "w", encoding='utf-8') as file: file.write(all_content)
内容的提问来源于stack exchange,提问作者Dhiraz Singh
相关产品推荐
相关产品推荐

