使用BeautifulSoup、urllib及Selenium爬取同级元素内容遇阻求助
Let’s break down practical solutions for each tool based on the page structure you shared:
1. BeautifulSoup + urllib
Since the answers are wrapped in <strong> tags right after their matching <span> questions, we can target them directly without overcomplicating things. Here’s a straightforward approach:
from urllib.request import urlopen from bs4 import BeautifulSoup # Fetch and parse the page page = urlopen("your-target-url") soup = BeautifulSoup(page, "html.parser") # First locate the p2 container holding all questions/answers p2_container = soup.find("p", class_="p2") # Get Answer1: Find the "Question1" span, then grab its next strong sibling question1_span = p2_container.find("span", string="Question1") answer1 = question1_span.find_next_sibling("strong").text.strip() # Get Answer2 using the same logic question2_span = p2_container.find("span", string="Question2") answer2 = question2_span.find_next_sibling("strong").text.strip() print(f"Answer1: {answer1}, Answer2: {answer2}")
If you’re sure the question order never changes, you can also skip targeting the spans and just grab all <strong> tags in the container:
answers = p2_container.find_all("strong") answer1 = answers[0].text.strip() answer2 = answers[1].text.strip()
2. Selenium
Your attempt to use the parent element might have failed because you didn’t wait for dynamic content to load, or you weren’t drilling into the parent’s children correctly. Try these reliable fixes:
Option 1: Directly target the answer with XPath
This skips the parent element step and links Answer1 directly to Question1 via XPath:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() driver.get("your-target-url") # Wait for the element to fully load (critical for dynamic pages) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//span[text()='Question1']/following-sibling::strong[1]")) ) # Grab the answers answer1 = driver.find_element(By.XPATH, "//span[text()='Question1']/following-sibling::strong[1]").text.strip() answer2 = driver.find_element(By.XPATH, "//span[text()='Question2']/following-sibling::strong[1]").text.strip() print(f"Answer1: {answer1}, Answer2: {answer2}") driver.quit()
Option 2: Use the parent p2 container correctly
If you still want to work with parent elements, make sure you select the container first then pull the strong tags inside:
# After loading the page and waiting for elements p2_container = driver.find_element(By.CLASS_NAME, "p2") answers = p2_container.find_elements(By.TAG_NAME, "strong") answer1 = answers[0].text.strip() answer2 = answers[1].text.strip()
The biggest mistake here is usually not waiting for the page to render—WebDriverWait ensures the elements exist before you try to scrape them.
内容的提问来源于stack exchange,提问作者Karma

