如何实现Selenium异步运行?加速Khan Academy爬虫脚本
Hey there! Let's tackle your Khan Academy scraping speed issues step by step. I've reviewed your script and have several actionable changes that should significantly boost performance. Here's what we'll adjust:
1. Reuse Selenium Drivers (Avoid Recreating Them)
Your current script initializes a new Chrome driver for each course and again for each profile's projects section. Starting and stopping browsers is one of the biggest performance hits. Instead, initialize a driver with optimizations once per task, or use a reusable driver setup.
Example Adjustment:
from selenium.webdriver.chrome.options import Options def init_driver(): chrome_options = Options() chrome_options.add_argument("--headless=new") # Faster, more compatible headless mode (Chrome 112+) chrome_options.add_argument("--disable-images") # Cut page load time by skipping images chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") # Fix memory issues in containers/VPS chrome_options.add_argument("--disable-extensions") # Disable unnecessary extensions driver = webdriver.Chrome(options=chrome_options) return driver, WebDriverWait(driver, 10) # Lower wait timeout from 15s to 10s (adjust if needed)
2. Ditch requests-html—Stick to Selenium for All Dynamic Content
You're mixing requests-html and Selenium, which means launching two separate headless browser instances. This redundant startup adds unnecessary latency. Use Selenium for all dynamic page interactions to keep things consistent and fast.
Example: Fetch Course Links with Selenium
def get_course_links(): driver, wait = init_driver() try: driver.get('https://www.khanacademy.org/computing/computer-programming/programming#intro-to-programming') wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'link_1uvuyao-o_O-nodeStyle_cu2reh-o_O-nodeStyleIcon_4udnki'))) soup = BeautifulSoup(driver.page_source, 'lxml') # Use lxml parser (faster than html.parser) courses_links = soup.find_all(class_='link_1uvuyao-o_O-nodeStyle_cu2reh-o_O-nodeStyleIcon_4udnki') list_courses = {} for links in courses_links: course = links.extract() link_course = course['href'] title_course = links.find(class_='nodeTitle_145jbuf') span_title_course = title_course.span text_span = span_title_course.text.strip() final_link_course = 'https://www.khanacademy.org' + link_course list_courses[text_span] = final_link_course return list_courses.values() finally: driver.quit()
3. Optimize "Show More" Click Logic
Your current loop uses presence_of_element_located, but we can make it more efficient by checking if the element is clickable (not just present) and using JavaScript clicks to avoid element intercept issues. Add a short delay between clicks to let content load without overwhelming the server.
Example: Improved Show More Click
def click_show_more(driver, wait): while True: try: showmore = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, 'button_1eqj1ga-o_O-shared_1t8r4tr-o_O-default_9fm203'))) driver.execute_script("arguments[0].click();", showmore) # JS click avoids "element not interactable" errors time.sleep(0.5) # Short pause to let content render except (TimeoutException, StaleElementReferenceException): break
4. Use Multiprocessing for Parallel Scraping
Since Selenium isn't thread-safe, multiprocessing is the way to go. Each process runs its own independent driver, letting you scrape multiple profiles at the same time. We'll use concurrent.futures.ProcessPoolExecutor for this.
Example: Parallel Profile Scraping
First, refactor your profile scraping into a standalone function:
from concurrent.futures import ProcessPoolExecutor def scrape_profile(profile_link): driver, wait = init_driver() try: print(f"Scraping {profile_link}") # Main profile page scrape driver.get(profile_link) soup = BeautifulSoup(driver.page_source, 'lxml') # Badge logic badge_list = soup.find_all(class_='badge-category-count') badgelist = [] if badge_list: for number in badge_list: text_num = number.text.strip() badgelist.append(text_num) number_badges = str(sum(map(int, badgelist))) # Handle cases where badge counts might be missing badge_challenge = badgelist[0] if len(badgelist)>=1 else 'NA' badge_lvl5 = badgelist[1] if len(badgelist)>=2 else 'NA' badge_lvl4 = badgelist[2] if len(badgelist)>=3 else 'NA' badge_lvl3 = badgelist[3] if len(badgelist)>=4 else 'NA' badge_lvl2 = badgelist[4] if len(badgelist)>=5 else 'NA' badge_lvl1 = badgelist[5] if len(badgelist)>=6 else 'NA' else: number_badges = badge_challenge = badge_lvl5 = badge_lvl4 = badge_lvl3 = badge_lvl2 = badge_lvl1 = 'NA' # User stats table user_info_table = soup.find('table', class_='user-statistics-table') if user_info_table: dates, points, videos = [tr.find_all('td')[1].text for tr in user_info_table.find_all('tr')] points = points.replace(",", "") # Clean up commas in numbers else: dates = points = videos = 'NA' # Discussion stats user_socio_table = soup.find_all('div', class_='discussion-stat') data = {} for gettext in user_socio_table: category = gettext.find('span') category_text = category.text.strip() number = category.previousSibling.strip() data[category_text] = number full_data_keys = ['questions','votes','answers','flags raised','project help requests','project help replies','comments','tips and thanks'] for header_value in full_data_keys: if header_value not in data: data[header_value] = 'NA' # Last activity date last_activity_date = 'NA' user_calendar = soup.find('div', class_='streak-calendar-scroll-container') if user_calendar: last_activity = user_calendar.find('span', class_='streak-cell filled') if last_activity: last_activity_date = last_activity.get('title', 'NA') # Top question votes driver.get(f"{profile_link}discussion/questions") soup = BeautifulSoup(driver.page_source, 'lxml') topq_votes = soup.find(class_='text_12zg6rl-o_O-LabelXSmall_mbug0d-o_O-votesSum_19las6u') topq_votes = re.findall(r'\d+', topq_votes.text.strip())[0] if topq_votes else '0' # Top answer votes driver.get(f"{profile_link}discussion/answers") soup = BeautifulSoup(driver.page_source, 'lxml') topa_votes = soup.find(class_='text_12zg6rl-o_O-LabelXSmall_mbug0d-o_O-votesSum_19las6u') topa_votes = re.findall(r'\d+', topa_votes.text.strip())[0] if topa_votes else '0' # Projects section driver.get(f"{profile_link}projects") click_show_more(driver, wait) soup = BeautifulSoup(driver.page_source, 'lxml') project_count = str(len(soup.find_all(class_='title_1usue9n'))) votes_spins = soup.find_all(class_='stats_35behe') total_votes = 0 total_spins = 0 for vs in votes_spins: parts = vs.text.strip().split() if len(parts) >=4: total_votes += int(parts[0]) spin_val = int(parts[3]) total_spins += spin_val if spin_val >=0 else 0 # Return a list matching CSV headers return [ profile_link, dates, points, videos, data['questions'], data['votes'], data['answers'], data['flags raised'], data['project help requests'], data['project help replies'], data['comments'], data['tips and thanks'], last_activity_date, project_count, str(total_votes), str(total_spins), topq_votes, topa_votes, number_badges, badge_lvl1, badge_lvl2, badge_lvl3, badge_lvl4, badge_lvl5, badge_challenge ] except Exception as e: print(f"Error scraping {profile_link}: {str(e)}") return [profile_link] + ['NA']*24 # Return NA for all fields if something breaks finally: driver.quit()
Then run the parallel scraping:
if __name__ == "__main__": # Step 1: Fetch all course links print("Fetching course links...") course_links = get_course_links() # Step 2: Collect all unique profile links print("Collecting profile links from courses...") all_profiles = [] for course in course_links: driver, wait = init_driver() try: driver.get(course) click_show_more(driver, wait) soup = BeautifulSoup(driver.page_source, 'lxml') profiles = soup.find_all(href=re.compile("/profile/kaid")) for links in profiles: text_link = links['href'] # Clean up links ending with /discussion text_link_clean = text_link[:-10] if text_link.endswith('/discussion') else text_link final_profile_link = 'https://www.khanacademy.org' + text_link_clean all_profiles.append(final_profile_link) print(f"Collected {len(profiles)} profiles from {course}") finally: driver.quit() # Remove duplicates all_profiles = list(set(all_profiles)) print(f"Total unique profiles to scrape: {len(all_profiles)}") # Step 3: Scrape profiles in parallel print("Starting parallel scraping...") with ProcessPoolExecutor(max_workers=4) as executor: # Adjust based on your CPU/RAM (start with 2-4) results = executor.map(scrape_profile, all_profiles) # Step 4: Write all results to CSV at once (faster than逐行写入) filename = "khan_withprojectandvotes.csv" headers = "link, date_joined, points, videos, questions, votes, answers, flags, project_request, project_replies, comments, tips_thx, last_date, number_project, projet_votes, projets_spins, topq_votes, topa_votes, sum_badges, badge_lvl1, badge_lvl2, badge_lvl3, badge_lvl4, badge_lvl5, badge_challenge\n" with open(filename, "w", encoding="utf-8") as f: f.write(headers) for row in results: # Clean up commas in data to avoid CSV formatting issues csv_row = ",".join(str(item).replace(",", "") for item in row) + "\n" f.write(csv_row) print(f"Scraping complete! Data saved to {filename}")
5. Additional Small Optimizations
- Use
lxmlParser: Install withpip install lxml—it's 2-3x faster than Python's built-inhtml.parser. - Batch CSV Writes: Collect all data first, then write once instead of after each profile—this cuts down on slow disk IO operations.
- Cache Scraped Profiles: If you run the script multiple times, save scraped links to a text file and skip them in future runs.
内容的提问来源于stack exchange,提问作者RobZ

