如何结合Selenium与Beautiful Soup爬取完整谷歌学术个人主页
Combine Selenium & BeautifulSoup to Scrape Full Google Scholar Profile
Hey there! To grab the complete list of articles from that Google Scholar profile, we need to first use Selenium to repeatedly click the "Show more" button until all content loads, then pass the fully loaded page to BeautifulSoup for data extraction. Here's a complete, working solution:
Step-by-Step Explanation & Code
First, make sure you have the required libraries installed (selenium, beautifulsoup4). Then use this combined script:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time # Target Google Scholar profile URL scholar_url = "https://scholar.google.com/citations?user=VjJm3zYAAAAJ&hl=en" # Initialize Chrome Driver (update the path to match your chromedriver location) driver = webdriver.Chrome(executable_path="/Applications/chromedriver84") driver.get(scholar_url) try: # Auto-click "Show more" until no more content is available while True: try: # Wait for the button to be clickable (not just present) show_more_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.ID, "gsc_bpf_more")) ) show_more_button.click() # Add a short delay to let new articles load time.sleep(2) except: # Break the loop when the button disappears (no more content) print("All content loaded!") break # Extract the fully loaded page source and parse with BeautifulSoup full_page_html = driver.page_source soup = BeautifulSoup(full_page_html, "html.parser") # Scrape article titles and authors articles = soup.find_all("tr", class_="gsc_a_tr") for count, article in enumerate(articles, 1): title = article.find("a", class_="gsc_a_at").get_text(strip=True) authors = article.find("div", class_="gs_gray").get_text(strip=True) print(f"Article {count}:") print(f"Title: {title}") print(f"Authors: {authors}\n") finally: # Always close the driver to free up resources driver.quit()
Key Improvements Over Your Original Code
- Repeated Button Clicks: Your original Selenium code only clicked "Show more" once—this script loops until the button is gone, ensuring all articles are loaded.
- Clickable Element Check: Uses
EC.element_to_be_clickableinstead ofpresence_of_element_locatedto avoid trying to click a button that's not ready yet. - Clean Data Extraction: Uses
get_text(strip=True)to remove extra whitespace from titles/authors for cleaner output. - Proper Resource Cleanup: The
finallyblock ensures the browser window closes even if an error occurs.
内容的提问来源于stack exchange,提问作者Noah
相关产品推荐
相关产品推荐

