Python网页爬取xlsx文件代码无法正常运行求助
爬取网页中.xlsx文件的问题修复
原代码无法正常运行,主要存在以下几个问题:
- 目标网站部分内容为动态加载,
requests.get仅能获取初始静态HTML,实际的.xlsx链接可能未包含在返回页面中 - 并非所有
<a>标签都带有href属性,直接访问link['href']会触发KeyError - 即使找到相对路径的链接,
pd.read_excel需要完整的绝对URL才能正常读取文件 - 网站可能拦截默认的
requests请求头,需添加合法的User-Agent模拟浏览器访问
修正后的代码
import requests from bs4 import BeautifulSoup import pandas as pd from urllib.parse import urljoin # 添加合法请求头,模拟浏览器访问 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } url = 'https://www.directliquidation.com/' try: response = requests.get(url, headers=headers) response.raise_for_status() # 检查请求是否成功 soup = BeautifulSoup(response.text, 'html.parser') excel_links = [] for link in soup.find_all('a'): # 先判断href属性是否存在,再检查文件后缀 href = link.get('href') if href and href.endswith('.xlsx'): # 将相对路径转换为绝对URL full_url = urljoin(url, href) excel_links.append(full_url) if not excel_links: print("未找到任何.xlsx文件链接") else: for idx, excel_link in enumerate(excel_links, 1): print(f"--- 第{idx}个Excel文件内容 ---") try: data = pd.read_excel(excel_link) print(data.head()) except Exception as e: print(f"读取文件失败: {str(e)}") except Exception as e: print(f"请求网页失败: {str(e)}")
动态内容补充方案
如果目标网站的.xlsx链接是通过JavaScript动态生成的,静态爬取方法无法获取,可使用selenium模拟浏览器加载动态内容:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup from urllib.parse import urljoin import pandas as pd options = Options() options.add_argument('--headless=new') # 无头模式运行浏览器 driver = webdriver.Chrome(options=options) driver.get('https://www.directliquidation.com/') # 等待页面加载完成(可根据实际情况调整等待时间) driver.implicitly_wait(10) soup = BeautifulSoup(driver.page_source, 'html.parser') excel_links = [] for link in soup.find_all('a'): href = link.get('href') if href and href.endswith('.xlsx'): full_url = urljoin('https://www.directliquidation.com/', href) excel_links.append(full_url) # 后续读取Excel文件的逻辑和之前一致 if excel_links: for idx, excel_link in enumerate(excel_links, 1): print(f"--- 第{idx}个Excel文件内容 ---") try: data = pd.read_excel(excel_link) print(data.head()) except Exception as e: print(f"读取文件失败: {str(e)}") else: print("未找到任何.xlsx文件链接") driver.quit()
内容的提问来源于stack exchange,提问作者Farid Porte
相关产品推荐
相关产品推荐

