You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Selenium的Session ID传递给Python Requests并实现单词抓取排序?

Hey there! No worries about being a Python newbie—we’ve all been in your shoes, and terminology slips are totally okay. Let’s break down how to build exactly what you’re asking for: using Selenium to refresh the random words, passing its session to Requests to grab the updated page, extracting the words, sorting them, and printing both the original and sorted lists.

Step-by-Step Explanation & Code

First, let’s make sure we cover all your requirements: we’ll use Selenium to trigger the refresh (since we need the updated words), transfer its session cookies to Requests so we fetch the same refreshed page, parse the HTML to get the words, then sort and print everything.

1. Install Required Packages (if you haven’t already)

You mentioned you have Requests, but we’ll need Selenium and BeautifulSoup for parsing. Run these in your terminal:

pip install selenium beautifulsoup4 webdriver-manager

We use webdriver-manager to auto-handle browser drivers (no need to manually download ChromeDriver/FirefoxDriver).

2. Full Working Code

Here’s the code with detailed comments to walk you through each part:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.service import Service as ChromeService
from webdriver_manager.chrome import ChromeDriverManager
import requests
from bs4 import BeautifulSoup

# Initialize Selenium Chrome driver (auto-installs the correct driver version)
driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()))

try:
    # Open the random words website
    driver.get("https://www.randomlists.com/random-words")

    # Wait for the "Refresh" button to load, then click it to generate new words
    # We use WebDriverWait to avoid errors if the page loads slowly
    refresh_button = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.ID, "refresh"))
    )
    refresh_button.click()

    # Short pause to let the new words load (adjust to 2s if needed)
    driver.implicitly_wait(1)

    # Extract cookies from Selenium's session and convert to Requests-compatible format
    # This ensures Requests fetches the exact refreshed page Selenium saw
    cookies = driver.get_cookies()
    requests_session = requests.Session()
    for cookie in cookies:
        requests_session.cookies.set(cookie['name'], cookie['value'])

finally:
    # Close the Selenium browser window—we don't need it anymore
    driver.quit()

# Use Requests with the transferred session to get the refreshed page
response = requests_session.get("https://www.randomlists.com/random-words")
response.raise_for_status()  # Trigger an error if the request fails (e.g., 404, 500)

# Parse the HTML to extract the random words
soup = BeautifulSoup(response.text, 'html.parser')
# The words live inside <li> elements under the <ul id="random_words"> container
word_elements = soup.select("#random_words li")
original_words = [element.text.strip() for element in word_elements]

# Sort the words alphabetically
sorted_words = sorted(original_words)

# Print the results in a readable format
print("Original Random Words:")
print("- " + "\n- ".join(original_words))
print("\nAlphabetically Sorted Words:")
print("- " + "\n- ".join(sorted_words))

3. Key Parts Explained

  • Session Transfer: We grab cookies from Selenium’s browser session and add them to a Requests Session object. This is critical because it ensures Requests fetches the exact refreshed page Selenium interacted with—without this, you might get the default initial words instead of the updated set.
  • Waiting for Elements: WebDriverWait makes sure we only click the refresh button once it’s fully loaded, which prevents "element not found" errors on slower connections.
  • HTML Parsing: Using CSS selectors (#random_words li) targets the exact elements holding the words, which is more reliable than scraping raw text from the page.

4. Quick Tips for Newbies

  • If you prefer Firefox over Chrome, replace ChromeService and ChromeDriverManager with FirefoxService and GeckoDriverManager (adjust imports accordingly).
  • If you still get old words after clicking refresh, increase the driver.implicitly_wait(1) value to 2 or 3 seconds—sometimes the page takes a bit longer to load new content.
  • If the refresh button’s ID changes (sites occasionally update their HTML), right-click the button in your browser, select "Inspect", and check the new ID to update the By.ID value.

内容的提问来源于stack exchange,提问作者kjolnir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:14:01