PythonAnywhere部署的Futbin网页爬虫无输出不写入文件问题求助
问题排查与解决方案
核心问题定位
- 第一,代码使用了**裸
except: pass**语句,所有运行时异常被直接吞掉,无法定位错误点,这是你看不到任何报错的核心原因 - 第二,PythonAnywhere免费账户有出站访问白名单限制,未在白名单内的站点会直接被拦截,同时Futbin有反爬策略,默认的requests请求头大概率会被屏蔽
- 第三,存储路径的父目录如果不存在,
open()的追加模式不会自动创建父目录,会直接抛出IO错误,被吞后就表现为无法写入文件 - 第四,Python定时任务默认开启输出缓冲,
print内容不会实时写入日志,若程序提前中断就看不到任何打印输出
修复步骤
1. 移除裸except,添加异常打印
把原代码里的except: pass替换为异常打印逻辑,方便在PythonAnywhere的定时任务日志中查看具体报错:
except Exception as e: print(f"运行出错: {repr(e)}", flush=True)
2. 处理访问限制与反爬
- 若使用PythonAnywhere免费账户,先确认Futbin是否在官方白名单内,不在的话需要升级付费账户解锁全量网络访问
- 给requests请求添加合法User-Agent头,模拟浏览器访问,避免被反爬拦截:
headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } page = requests.get(URL, cookies=cookies, headers=headers)
3. 修复路径写入问题
提前创建存储目录,避免父目录不存在导致的IO错误:
首先导入os模块,然后在写文件前添加目录创建逻辑:
import os # 在open语句前添加 os.makedirs("/home/exec85/scrape", exist_ok=True) with open("/home/exec85/scrape/pc.txt", "a") as f: f.write(f"{list_all_results}\n") # 建议加换行符区分每次运行的内容
4. 关闭输出缓冲
两种方案二选一即可:
- 所有
print语句添加flush=True参数,例如print("Scraping page " + str(i) + "/745", flush=True) - 在PythonAnywhere定时任务的执行命令中添加
-u参数,例如python3 -u /home/exec85/scrape/你的脚本文件名.py
修改后完整参考代码
import requests from bs4 import BeautifulSoup import time import random import os list_all_results = [] # 提前创建存储目录 os.makedirs("/home/exec85/scrape", exist_ok=True) # 统一请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } for i in range(1, 3): time.sleep(random.uniform(1.5, 2)) print("Scraping page " + str(i) + "/745", flush=True) try: URL = "https://www.futbin.com/players?page=" + str(i) platform = "pc" cookies = {"platform": platform} page = requests.get(URL, cookies=cookies, headers=headers, timeout=10) # 可选:添加状态码校验 page.raise_for_status() soup = BeautifulSoup(page.content, "html.parser") result_names = soup.find_all("a", attrs={"class": "player_name_players_table"}) result_ratings = soup.find_all( "span", attrs={"class": lambda r: r.startswith("form rating ut21") if r else False}, ) result_rarity = soup.find_all("td", {"class": "mobile-hide-table-col"}) result_prices_pc = soup.find_all( "span", attrs={"class": "pc_color font-weight-bold"} ) list_names = [] list_ratings = [] list_rarities = [] list_prices = [] for name in result_names: list_names.append(name.text) for rating in result_ratings: list_ratings.append(rating.text) for rarity in result_rarity: list_rarities.append(rarity.text) for price in result_prices_pc: n = price.text.strip() if "K" in n: n2 = n.replace("K", "") full_int = int(float(n2) * 1000) list_prices.append(full_int) elif "M" in n: n2 = n.replace("M", "") full_int = int(float(n2) * 1000000) list_prices.append(full_int) else: list_prices.append(int(price.text.strip())) int_list_length = len(list_names) for idx in range(0, int_list_length): list_all_results.append( tuple( (list_names[idx], list_ratings[idx], list_rarities[idx], list_prices[idx]) ) ) with open("/home/exec85/scrape/pc.txt", "a") as f: f.write(f"{list_all_results}\n") except Exception as e: print(f"第{i}页抓取出错: {repr(e)}", flush=True) print("FINISHED", flush=True)
内容的提问来源于stack exchange,提问作者exec85
相关产品推荐
相关产品推荐

