如何解决FastAPI爬虫API数据仅在控制台显示、无法在浏览器正常展示的问题?
解决FastAPI爬虫API数据无法在浏览器展示的问题
我帮你排查了代码里的几个关键问题,咱们一步步来修复:
核心问题:爬虫方法没有返回数据
你的scrapedata方法里只打印了DataFrame,但没有任何return语句,所以FastAPI的接口函数拿不到要返回的数据,浏览器自然显示不了内容。这是最关键的问题!
其他需要修正的细节
除了返回数据,还有几个地方会影响功能或性能:
1. FastAPI接口的同步/异步问题
Selenium是同步阻塞的操作,你的接口用了async定义,但同步代码在异步函数里可能会导致阻塞整个服务。建议把接口改成普通的同步函数(去掉async)。
2. 爬虫初始化时机错误
你在FastAPI代码的全局作用域里调用了data = scrape()和data.scrapedata(),这会导致服务器启动时就自动执行一次爬虫,而不是用户请求/data接口时才执行。应该把爬虫的实例化和调用放到接口函数内部。
3. Windows路径的转义问题
ChromeDriver的路径用了"C:\Program Files (x86)\chromedriver.exe",这里的反斜杠会被当成转义字符,导致路径识别错误。应该改成原始字符串(加r前缀):r"C:\Program Files (x86)\chromedriver.exe"。
4. 硬编码循环次数的风险
你用range(1, 35)硬编码要爬取的歌曲数量,如果页面上的歌曲数量不足35,会直接抛出元素找不到的错误。建议先获取所有歌曲元素,再遍历它们。
修改后的完整代码
修正后的Scraper代码
import time from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException import pandas as pd class Scrape: # 类名采用大驼峰命名规范更符合Python惯例 def scrapedata(self): # 用原始字符串修正Windows路径转义问题 ser = Service(r"C:\Program Files (x86)\chromedriver.exe") options = webdriver.ChromeOptions() options.add_experimental_option('excludeSwitches', ['enable-logging']) # 添加无头模式,避免每次请求弹出浏览器窗口 options.add_argument('--headless=new') driver = webdriver.Chrome(options=options,service=ser) driver.get('https://soundcloud.com/jujubucks') print(driver.title) wait = WebDriverWait(driver,30) wait.until(EC.element_to_be_clickable((By.ID,"onetrust-accept-btn-handler"))).click() song_list = [] # 先获取所有歌曲元素再遍历,避免硬编码数量导致的报错 wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "soundList__item"))) song_contents = driver.find_elements(By.CLASS_NAME, "soundList__item") for item in song_contents: driver.execute_script("arguments[0].scrollIntoView(true);", item) try: search = item.find_element(By.XPATH, ".//a[contains(@class,'soundTitle__username')]/span").text search_song = item.find_element(By.XPATH, ".//a[contains(@class,'soundTitle__title')]/span").text search_date = item.find_element(By.XPATH, ".//time[contains(@class,'relativeTime')]/span").text search_plays = item.find_element(By.XPATH, ".//span[contains(@class,'sc-ministats-small')]/span").text except NoSuchElementException: continue # 修正判断逻辑:search_plays是字符串,不会等于布尔值False if not search_plays: continue option ={ 'Artist': search, 'Song_title': search_song, 'Date': search_date, 'Streams': search_plays } song_list.append(option) df = pd.DataFrame(song_list) print(df) driver.quit() # 返回FastAPI可直接序列化的字典列表 return df.to_dict('records')
修正后的FastAPI代码
from fastapi import FastAPI from Scraper import Scrape # 对应修改后的类名 app = FastAPI() # 改成同步函数,将爬虫逻辑移至接口内部实现按需调用 @app.get("/data") def get_songs(): scraper = Scrape() return scraper.scrapedata()
验证步骤
- 替换上述代码后,重启Uvicorn服务器
- 访问
http://localhost:8000/data(或你配置的自定义端口) - 现在浏览器应该能正常展示爬取到的歌曲数据列表了
内容的提问来源于stack exchange,提问作者Houston Khanyile
相关产品推荐
相关产品推荐

