You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup爬虫无法获取信息 适配Python3后导出CSV为空

问题定位
  • 缺失os模块导入:代码中调用了os.path.exists、os.makedirs、os.path.join等方法,但开头未导入os模块,部分环境会直接终止执行。
  • 表格筛选逻辑写反:核心问题,世界杯参赛阵容的表格同时携带sortable和wikitable两个类属性,原代码写的if "wikitable" not in table['class']会直接跳过所有有效阵容表,导致最终收集的结果为空数组,导出的CSV自然为空。
  • URL协议与年份切片错误:维基百科已全站强制HTTPS,原代码用HTTP协议请求会被重定向,且HTTP/HTTPS的URL长度不同,原固定切片url[29:33]无法正确提取年份。
  • 缓存文件读取未指定编码:打开本地HTML缓存时未指定utf-8编码,Windows环境下会默认用GBK编码读取,导致内容乱码无法匹配到页面元素。
修复方案

对应问题逐一调整即可,修复后的完整代码如下:

import os
import hashlib
import requests
from bs4 import BeautifulSoup
import pandas as pd

if not os.path.exists('.cache'):
    os.makedirs('.cache')

ua = 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/34.0.1847.116 Safari/15612.1.29.41.4'
session = requests.Session()

def get(url):
    '''Return cached lxml tree for url'''
    path = os.path.join('.cache', hashlib.sha256(url.encode('utf-8')).hexdigest() + '.html')
    if not os.path.exists(path):
        print(url)
        response = session.get(url, headers={'User-Agent': ua})
        with open(path, 'wb') as fd:
            fd.write(response.text.encode('utf-8'))
    # 读取时指定utf-8编码
    return BeautifulSoup(open(path, encoding='utf-8'), 'html.parser')

def squads(url):
    result = []
    soup = get(url)
    # HTTPS协议下切片位置调整为30:34
    year = url[30:34]
    for table in soup.find_all('table','sortable'):
        # 筛选逻辑修正为包含wikitable才处理
        if "wikitable" in table['class']:
            country = table.find_previous("span","mw-headline").text
            for tr in table.find_all('tr')[1:]:
                cells = [td.text.strip() for td in tr.find_all('td')]
                # 空行跳过避免报错
                if not cells:
                    continue
                cells += [country, cells[1].a.get('title') if cells[1].a else 'none', year]
                result.append(cells)
    return result

years = list(range(1930,1939,4)) + list(range(1950,2019,4))
result = []
for year in years:
    # 协议改为HTTPS
    url = "https://en.wikipedia.org/wiki/"+str(year)+"_FIFA_World_Cup_squads"
    result += squads(url)

Final_result = pd.DataFrame(result)
# 编码改为utf-8-sig兼容Excel打开,可根据需求调整
Final_result.to_csv('/Users/home/Downloads/data.csv', index=False, encoding='utf-8-sig')

内容的提问来源于stack exchange,提问作者Juan Carlos Saravia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 15:18:03