You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup提取YC公司页面标题无结果,求解决方法

修复方案:提取Y Combinator公司标题

问题原因

原代码无法获取结果,是因为该页面采用JavaScript动态渲染——公司数据并非直接包含在初始HTML中,而是页面加载后通过JS异步请求加载的。requests.get()只能获取静态HTML内容,无法获取JS动态生成的元素,因此BeautifulSoup找不到目标标签。

解决方案1:使用Selenium模拟浏览器加载

Selenium可以模拟真实浏览器的行为,等待页面完全加载后再提取数据。

步骤:

  1. 安装依赖:
pip install selenium
  1. 下载对应浏览器的驱动(如ChromeDriver,需与浏览器版本匹配),并确保驱动路径可被Python访问。

修复后的代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

url = 'https://www.ycombinator.com/companies?industry=B2B%20Software%20and%20Services&status=Active&status=Public&status=Inactive&tags=Fintech&tags=Developer%20Tools&tags=Artificial%20Intelligence&tags=Analytics'

# 初始化Chrome浏览器(需确保ChromeDriver路径正确)
driver = webdriver.Chrome()
driver.get(url)

# 等待页面加载(可根据网络情况调整等待时间)
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CLASS_NAME, 'company-name'))
)

# 滚动页面加载更多内容(如果需要获取全部公司)
last_height = driver.execute_script("return document.body.scrollHeight")
while True:
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)
    new_height = driver.execute_script("return document.body.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

# 获取页面源码并解析
soup = BeautifulSoup(driver.page_source, 'html.parser')
driver.quit()

# 提取公司标题
company_names = soup.find_all('span', class_='company-name')
for name in company_names:
    print(name.get_text(strip=True))

解决方案2:直接调用官方API(更高效)

通过浏览器开发者工具的"网络"面板,可以找到页面加载公司数据的API接口,直接请求该接口获取JSON数据,无需模拟浏览器。

修复后的代码:

import requests

# 抓包得到的API接口(参数可根据需求调整)
api_url = 'https://www.ycombinator.com/companies/export.json?industry=B2B%20Software%20and%20Services&status=Active&status=Public&status=Inactive&tags=Fintech&tags=Developer%20Tools&tags=Artificial%20Intelligence&tags=Analytics'

response = requests.get(api_url)
data = response.json()

# 提取公司名称
for company in data:
    print(company['name'])

说明:

这个API返回的是结构化JSON数据,包含公司名称、状态、标签等所有信息,比解析HTML更高效可靠。如果后续API路径有变化,可通过浏览器开发者工具重新抓取。

内容的提问来源于stack exchange,提问作者puru singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 15:53:12