You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取xlsx文件代码无法正常运行求助

爬取网页中.xlsx文件的问题修复

原代码无法正常运行,主要存在以下几个问题:

  • 目标网站部分内容为动态加载,requests.get仅能获取初始静态HTML,实际的.xlsx链接可能未包含在返回页面中
  • 并非所有<a>标签都带有href属性,直接访问link['href']会触发KeyError
  • 即使找到相对路径的链接,pd.read_excel需要完整的绝对URL才能正常读取文件
  • 网站可能拦截默认的requests请求头,需添加合法的User-Agent模拟浏览器访问

修正后的代码

import requests
from bs4 import BeautifulSoup
import pandas as pd
from urllib.parse import urljoin

# 添加合法请求头,模拟浏览器访问
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

url = 'https://www.directliquidation.com/'

try:
    response = requests.get(url, headers=headers)
    response.raise_for_status()  # 检查请求是否成功
    soup = BeautifulSoup(response.text, 'html.parser')

    excel_links = []
    for link in soup.find_all('a'):
        # 先判断href属性是否存在,再检查文件后缀
        href = link.get('href')
        if href and href.endswith('.xlsx'):
            # 将相对路径转换为绝对URL
            full_url = urljoin(url, href)
            excel_links.append(full_url)

    if not excel_links:
        print("未找到任何.xlsx文件链接")
    else:
        for idx, excel_link in enumerate(excel_links, 1):
            print(f"--- 第{idx}个Excel文件内容 ---")
            try:
                data = pd.read_excel(excel_link)
                print(data.head())
            except Exception as e:
                print(f"读取文件失败: {str(e)}")

except Exception as e:
    print(f"请求网页失败: {str(e)}")

动态内容补充方案

如果目标网站的.xlsx链接是通过JavaScript动态生成的,静态爬取方法无法获取,可使用selenium模拟浏览器加载动态内容:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
from urllib.parse import urljoin
import pandas as pd

options = Options()
options.add_argument('--headless=new')  # 无头模式运行浏览器
driver = webdriver.Chrome(options=options)
driver.get('https://www.directliquidation.com/')

# 等待页面加载完成(可根据实际情况调整等待时间)
driver.implicitly_wait(10)

soup = BeautifulSoup(driver.page_source, 'html.parser')
excel_links = []
for link in soup.find_all('a'):
    href = link.get('href')
    if href and href.endswith('.xlsx'):
        full_url = urljoin('https://www.directliquidation.com/', href)
        excel_links.append(full_url)

# 后续读取Excel文件的逻辑和之前一致
if excel_links:
    for idx, excel_link in enumerate(excel_links, 1):
        print(f"--- 第{idx}个Excel文件内容 ---")
        try:
            data = pd.read_excel(excel_link)
            print(data.head())
        except Exception as e:
            print(f"读取文件失败: {str(e)}")
else:
    print("未找到任何.xlsx文件链接")

driver.quit()

内容的提问来源于stack exchange,提问作者Farid Porte

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 20:47:13