You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何结合Selenium与Beautiful Soup爬取完整谷歌学术个人主页

Combine Selenium & BeautifulSoup to Scrape Full Google Scholar Profile

Hey there! To grab the complete list of articles from that Google Scholar profile, we need to first use Selenium to repeatedly click the "Show more" button until all content loads, then pass the fully loaded page to BeautifulSoup for data extraction. Here's a complete, working solution:

Step-by-Step Explanation & Code

First, make sure you have the required libraries installed (selenium, beautifulsoup4). Then use this combined script:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

# Target Google Scholar profile URL
scholar_url = "https://scholar.google.com/citations?user=VjJm3zYAAAAJ&hl=en"

# Initialize Chrome Driver (update the path to match your chromedriver location)
driver = webdriver.Chrome(executable_path="/Applications/chromedriver84")
driver.get(scholar_url)

try:
    # Auto-click "Show more" until no more content is available
    while True:
        try:
            # Wait for the button to be clickable (not just present)
            show_more_button = WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.ID, "gsc_bpf_more"))
            )
            show_more_button.click()
            # Add a short delay to let new articles load
            time.sleep(2)
        except:
            # Break the loop when the button disappears (no more content)
            print("All content loaded!")
            break

    # Extract the fully loaded page source and parse with BeautifulSoup
    full_page_html = driver.page_source
    soup = BeautifulSoup(full_page_html, "html.parser")

    # Scrape article titles and authors
    articles = soup.find_all("tr", class_="gsc_a_tr")
    for count, article in enumerate(articles, 1):
        title = article.find("a", class_="gsc_a_at").get_text(strip=True)
        authors = article.find("div", class_="gs_gray").get_text(strip=True)
        print(f"Article {count}:")
        print(f"Title: {title}")
        print(f"Authors: {authors}\n")

finally:
    # Always close the driver to free up resources
    driver.quit()

Key Improvements Over Your Original Code

  • Repeated Button Clicks: Your original Selenium code only clicked "Show more" once—this script loops until the button is gone, ensuring all articles are loaded.
  • Clickable Element Check: Uses EC.element_to_be_clickable instead of presence_of_element_located to avoid trying to click a button that's not ready yet.
  • Clean Data Extraction: Uses get_text(strip=True) to remove extra whitespace from titles/authors for cleaner output.
  • Proper Resource Cleanup: The finally block ensures the browser window closes even if an error occurs.

内容的提问来源于stack exchange,提问作者Noah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 19:42:55