You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PythonAnywhere部署的Futbin网页爬虫无输出不写入文件问题求助

问题排查与解决方案

核心问题定位

  • 第一,代码使用了**裸except: pass**语句,所有运行时异常被直接吞掉,无法定位错误点,这是你看不到任何报错的核心原因
  • 第二,PythonAnywhere免费账户有出站访问白名单限制,未在白名单内的站点会直接被拦截,同时Futbin有反爬策略,默认的requests请求头大概率会被屏蔽
  • 第三,存储路径的父目录如果不存在,open()的追加模式不会自动创建父目录,会直接抛出IO错误,被吞后就表现为无法写入文件
  • 第四,Python定时任务默认开启输出缓冲,print内容不会实时写入日志,若程序提前中断就看不到任何打印输出

修复步骤

1. 移除裸except,添加异常打印

把原代码里的except: pass替换为异常打印逻辑,方便在PythonAnywhere的定时任务日志中查看具体报错:

except Exception as e:
    print(f"运行出错: {repr(e)}", flush=True)

2. 处理访问限制与反爬

  • 若使用PythonAnywhere免费账户,先确认Futbin是否在官方白名单内,不在的话需要升级付费账户解锁全量网络访问
  • 给requests请求添加合法User-Agent头,模拟浏览器访问,避免被反爬拦截:
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}
page = requests.get(URL, cookies=cookies, headers=headers)

3. 修复路径写入问题

提前创建存储目录,避免父目录不存在导致的IO错误:
首先导入os模块,然后在写文件前添加目录创建逻辑:

import os
# 在open语句前添加
os.makedirs("/home/exec85/scrape", exist_ok=True)
with open("/home/exec85/scrape/pc.txt", "a") as f:
    f.write(f"{list_all_results}\n") # 建议加换行符区分每次运行的内容

4. 关闭输出缓冲

两种方案二选一即可:

  • 所有print语句添加flush=True参数,例如print("Scraping page " + str(i) + "/745", flush=True)
  • 在PythonAnywhere定时任务的执行命令中添加-u参数,例如python3 -u /home/exec85/scrape/你的脚本文件名.py

修改后完整参考代码

import requests
from bs4 import BeautifulSoup
import time
import random
import os

list_all_results = []
# 提前创建存储目录
os.makedirs("/home/exec85/scrape", exist_ok=True)
# 统一请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

for i in range(1, 3):
    time.sleep(random.uniform(1.5, 2))
    print("Scraping page " + str(i) + "/745", flush=True)

    try:
        URL = "https://www.futbin.com/players?page=" + str(i)
        platform = "pc"
        cookies = {"platform": platform}
        page = requests.get(URL, cookies=cookies, headers=headers, timeout=10)
        # 可选:添加状态码校验
        page.raise_for_status()

        soup = BeautifulSoup(page.content, "html.parser")

        result_names = soup.find_all("a", attrs={"class": "player_name_players_table"})
        result_ratings = soup.find_all(
            "span",
            attrs={"class": lambda r: r.startswith("form rating ut21") if r else False},
        )
        result_rarity = soup.find_all("td", {"class": "mobile-hide-table-col"})
        result_prices_pc = soup.find_all(
            "span", attrs={"class": "pc_color font-weight-bold"}
        )

        list_names = []
        list_ratings = []
        list_rarities = []
        list_prices = []

        for name in result_names:
            list_names.append(name.text)

        for rating in result_ratings:
            list_ratings.append(rating.text)

        for rarity in result_rarity:
            list_rarities.append(rarity.text)

        for price in result_prices_pc:
            n = price.text.strip()
            if "K" in n:
                n2 = n.replace("K", "")
                full_int = int(float(n2) * 1000)
                list_prices.append(full_int)
            elif "M" in n:
                n2 = n.replace("M", "")
                full_int = int(float(n2) * 1000000)
                list_prices.append(full_int)
            else:
                list_prices.append(int(price.text.strip()))

        int_list_length = len(list_names)
        for idx in range(0, int_list_length):
            list_all_results.append(
                tuple(
                    (list_names[idx], list_ratings[idx], list_rarities[idx], list_prices[idx])
                )
            )

        with open("/home/exec85/scrape/pc.txt", "a") as f:
            f.write(f"{list_all_results}\n")

    except Exception as e:
        print(f"第{i}页抓取出错: {repr(e)}", flush=True)

print("FINISHED", flush=True)

内容的提问来源于stack exchange,提问作者exec85

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 00:12:03