You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现Selenium异步运行?加速Khan Academy爬虫脚本

Optimizations for Your Khan Academy Scraping Script

Hey there! Let's tackle your Khan Academy scraping speed issues step by step. I've reviewed your script and have several actionable changes that should significantly boost performance. Here's what we'll adjust:


1. Reuse Selenium Drivers (Avoid Recreating Them)

Your current script initializes a new Chrome driver for each course and again for each profile's projects section. Starting and stopping browsers is one of the biggest performance hits. Instead, initialize a driver with optimizations once per task, or use a reusable driver setup.

Example Adjustment:

from selenium.webdriver.chrome.options import Options

def init_driver():
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")  # Faster, more compatible headless mode (Chrome 112+)
    chrome_options.add_argument("--disable-images")  # Cut page load time by skipping images
    chrome_options.add_argument("--no-sandbox")
    chrome_options.add_argument("--disable-dev-shm-usage")  # Fix memory issues in containers/VPS
    chrome_options.add_argument("--disable-extensions")  # Disable unnecessary extensions
    driver = webdriver.Chrome(options=chrome_options)
    return driver, WebDriverWait(driver, 10)  # Lower wait timeout from 15s to 10s (adjust if needed)

2. Ditch requests-html—Stick to Selenium for All Dynamic Content

You're mixing requests-html and Selenium, which means launching two separate headless browser instances. This redundant startup adds unnecessary latency. Use Selenium for all dynamic page interactions to keep things consistent and fast.

def get_course_links():
    driver, wait = init_driver()
    try:
        driver.get('https://www.khanacademy.org/computing/computer-programming/programming#intro-to-programming')
        wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'link_1uvuyao-o_O-nodeStyle_cu2reh-o_O-nodeStyleIcon_4udnki')))
        soup = BeautifulSoup(driver.page_source, 'lxml')  # Use lxml parser (faster than html.parser)
        courses_links = soup.find_all(class_='link_1uvuyao-o_O-nodeStyle_cu2reh-o_O-nodeStyleIcon_4udnki')
        
        list_courses = {}
        for links in courses_links:
            course = links.extract()
            link_course = course['href']
            title_course = links.find(class_='nodeTitle_145jbuf')
            span_title_course = title_course.span
            text_span = span_title_course.text.strip()
            final_link_course = 'https://www.khanacademy.org' + link_course
            list_courses[text_span] = final_link_course
        return list_courses.values()
    finally:
        driver.quit()

3. Optimize "Show More" Click Logic

Your current loop uses presence_of_element_located, but we can make it more efficient by checking if the element is clickable (not just present) and using JavaScript clicks to avoid element intercept issues. Add a short delay between clicks to let content load without overwhelming the server.

Example: Improved Show More Click

def click_show_more(driver, wait):
    while True:
        try:
            showmore = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, 'button_1eqj1ga-o_O-shared_1t8r4tr-o_O-default_9fm203')))
            driver.execute_script("arguments[0].click();", showmore)  # JS click avoids "element not interactable" errors
            time.sleep(0.5)  # Short pause to let content render
        except (TimeoutException, StaleElementReferenceException):
            break

4. Use Multiprocessing for Parallel Scraping

Since Selenium isn't thread-safe, multiprocessing is the way to go. Each process runs its own independent driver, letting you scrape multiple profiles at the same time. We'll use concurrent.futures.ProcessPoolExecutor for this.

Example: Parallel Profile Scraping

First, refactor your profile scraping into a standalone function:

from concurrent.futures import ProcessPoolExecutor

def scrape_profile(profile_link):
    driver, wait = init_driver()
    try:
        print(f"Scraping {profile_link}")
        # Main profile page scrape
        driver.get(profile_link)
        soup = BeautifulSoup(driver.page_source, 'lxml')
        
        # Badge logic
        badge_list = soup.find_all(class_='badge-category-count')
        badgelist = []
        if badge_list:
            for number in badge_list:
                text_num = number.text.strip()
                badgelist.append(text_num)
            number_badges = str(sum(map(int, badgelist)))
            # Handle cases where badge counts might be missing
            badge_challenge = badgelist[0] if len(badgelist)>=1 else 'NA'
            badge_lvl5 = badgelist[1] if len(badgelist)>=2 else 'NA'
            badge_lvl4 = badgelist[2] if len(badgelist)>=3 else 'NA'
            badge_lvl3 = badgelist[3] if len(badgelist)>=4 else 'NA'
            badge_lvl2 = badgelist[4] if len(badgelist)>=5 else 'NA'
            badge_lvl1 = badgelist[5] if len(badgelist)>=6 else 'NA'
        else:
            number_badges = badge_challenge = badge_lvl5 = badge_lvl4 = badge_lvl3 = badge_lvl2 = badge_lvl1 = 'NA'
        
        # User stats table
        user_info_table = soup.find('table', class_='user-statistics-table')
        if user_info_table:
            dates, points, videos = [tr.find_all('td')[1].text for tr in user_info_table.find_all('tr')]
            points = points.replace(",", "")  # Clean up commas in numbers
        else:
            dates = points = videos = 'NA'
        
        # Discussion stats
        user_socio_table = soup.find_all('div', class_='discussion-stat')
        data = {}
        for gettext in user_socio_table:
            category = gettext.find('span')
            category_text = category.text.strip()
            number = category.previousSibling.strip()
            data[category_text] = number
        
        full_data_keys = ['questions','votes','answers','flags raised','project help requests','project help replies','comments','tips and thanks']
        for header_value in full_data_keys:
            if header_value not in data:
                data[header_value] = 'NA'
        
        # Last activity date
        last_activity_date = 'NA'
        user_calendar = soup.find('div', class_='streak-calendar-scroll-container')
        if user_calendar:
            last_activity = user_calendar.find('span', class_='streak-cell filled')
            if last_activity:
                last_activity_date = last_activity.get('title', 'NA')
        
        # Top question votes
        driver.get(f"{profile_link}discussion/questions")
        soup = BeautifulSoup(driver.page_source, 'lxml')
        topq_votes = soup.find(class_='text_12zg6rl-o_O-LabelXSmall_mbug0d-o_O-votesSum_19las6u')
        topq_votes = re.findall(r'\d+', topq_votes.text.strip())[0] if topq_votes else '0'
        
        # Top answer votes
        driver.get(f"{profile_link}discussion/answers")
        soup = BeautifulSoup(driver.page_source, 'lxml')
        topa_votes = soup.find(class_='text_12zg6rl-o_O-LabelXSmall_mbug0d-o_O-votesSum_19las6u')
        topa_votes = re.findall(r'\d+', topa_votes.text.strip())[0] if topa_votes else '0'
        
        # Projects section
        driver.get(f"{profile_link}projects")
        click_show_more(driver, wait)
        soup = BeautifulSoup(driver.page_source, 'lxml')
        
        project_count = str(len(soup.find_all(class_='title_1usue9n')))
        votes_spins = soup.find_all(class_='stats_35behe')
        
        total_votes = 0
        total_spins = 0
        for vs in votes_spins:
            parts = vs.text.strip().split()
            if len(parts) >=4:
                total_votes += int(parts[0])
                spin_val = int(parts[3])
                total_spins += spin_val if spin_val >=0 else 0
        
        # Return a list matching CSV headers
        return [
            profile_link, dates, points, videos,
            data['questions'], data['votes'], data['answers'], data['flags raised'],
            data['project help requests'], data['project help replies'], data['comments'],
            data['tips and thanks'], last_activity_date, project_count, str(total_votes),
            str(total_spins), topq_votes, topa_votes, number_badges, badge_lvl1, badge_lvl2,
            badge_lvl3, badge_lvl4, badge_lvl5, badge_challenge
        ]
    except Exception as e:
        print(f"Error scraping {profile_link}: {str(e)}")
        return [profile_link] + ['NA']*24  # Return NA for all fields if something breaks
    finally:
        driver.quit()

Then run the parallel scraping:

if __name__ == "__main__":
    # Step 1: Fetch all course links
    print("Fetching course links...")
    course_links = get_course_links()
    
    # Step 2: Collect all unique profile links
    print("Collecting profile links from courses...")
    all_profiles = []
    for course in course_links:
        driver, wait = init_driver()
        try:
            driver.get(course)
            click_show_more(driver, wait)
            soup = BeautifulSoup(driver.page_source, 'lxml')
            profiles = soup.find_all(href=re.compile("/profile/kaid"))
            for links in profiles:
                text_link = links['href']
                # Clean up links ending with /discussion
                text_link_clean = text_link[:-10] if text_link.endswith('/discussion') else text_link
                final_profile_link = 'https://www.khanacademy.org' + text_link_clean
                all_profiles.append(final_profile_link)
            print(f"Collected {len(profiles)} profiles from {course}")
        finally:
            driver.quit()
    
    # Remove duplicates
    all_profiles = list(set(all_profiles))
    print(f"Total unique profiles to scrape: {len(all_profiles)}")
    
    # Step 3: Scrape profiles in parallel
    print("Starting parallel scraping...")
    with ProcessPoolExecutor(max_workers=4) as executor:  # Adjust based on your CPU/RAM (start with 2-4)
        results = executor.map(scrape_profile, all_profiles)
    
    # Step 4: Write all results to CSV at once (faster than逐行写入)
    filename = "khan_withprojectandvotes.csv"
    headers = "link, date_joined, points, videos, questions, votes, answers, flags, project_request, project_replies, comments, tips_thx, last_date, number_project, projet_votes, projets_spins, topq_votes, topa_votes, sum_badges, badge_lvl1, badge_lvl2, badge_lvl3, badge_lvl4, badge_lvl5, badge_challenge\n"
    
    with open(filename, "w", encoding="utf-8") as f:
        f.write(headers)
        for row in results:
            # Clean up commas in data to avoid CSV formatting issues
            csv_row = ",".join(str(item).replace(",", "") for item in row) + "\n"
            f.write(csv_row)
    
    print(f"Scraping complete! Data saved to {filename}")

5. Additional Small Optimizations

  • Use lxml Parser: Install with pip install lxml—it's 2-3x faster than Python's built-in html.parser.
  • Batch CSV Writes: Collect all data first, then write once instead of after each profile—this cuts down on slow disk IO operations.
  • Cache Scraped Profiles: If you run the script multiple times, save scraped links to a text file and skip them in future runs.

内容的提问来源于stack exchange,提问作者RobZ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:12:29