如何用Python 3.7与Selenium遍历类名含'comment t1_'的<div>元素?
Hey there! Let’s tackle this Reddit bot challenge together. Since the find_element_by_class_name method is deprecated and those messy, random XPaths aren’t reliable, CSS selectors are your best bet here—they’re perfect for targeting elements where the class starts with a specific string.
Step 1: Grab All Target Comments with CSS Selectors
Reddit’s comment elements have classes starting with comment t1_, so we can use the CSS ^= attribute selector to match any element whose class attribute begins with that exact string. This is way cleaner and more stable than those unwieldy XPaths.
In modern Selenium, you’ll use the By class to define the selector type. Here’s how to fetch all relevant comment elements:
from selenium.webdriver.common.by import By # Get every comment element with a class starting with "comment t1_" comments = driver.find_elements(By.CSS_SELECTOR, '[class^="comment t1_"]')
Step 2: Parse Vote Counts from Comments
Next, you need to extract the vote count from each comment. Reddit usually displays scores in an element with a class containing score (like score or score-hidden). Adjust the selector below if Reddit updates its UI, but this should work for most cases:
def parse_score(score_text): """Convert Reddit's score text (e.g., "1.2k", "5", "•") to an integer""" score_text = score_text.strip().lower() if 'k' in score_text: return int(float(score_text.replace('k', '')) * 1000) elif 'm' in score_text: return int(float(score_text.replace('m', '')) * 1000000) else: # Handle hidden scores (displayed as "•") or plain numbers try: return int(score_text) except ValueError: return 0 # Skip hidden/invalid scores
Step 3: Track the Top-Voted Comment
Loop through all comments, parse their scores, and keep track of the comment with the highest vote count:
max_score = -1 top_comment_element = None for comment in comments: # Locate the score element inside the current comment score_elem = comment.find_element(By.CSS_SELECTOR, '[class*="score"]') score = parse_score(score_elem.text) if score > max_score: max_score = score top_comment_element = comment # Retrieve and print the top comment's content (adjust selector if needed) if top_comment_element: comment_content = top_comment_element.find_element(By.CSS_SELECTOR, '[data-testid="comment"]').text print(f"Top Comment (Score: {max_score}):\n{comment_content}") else: print("No valid comments found.")
Key Tips for Reliability
- Dynamic Loading: If comments load as you scroll, add a loop to scroll the page and wait for new elements to load. Use
WebDriverWaitto avoid race conditions (example included below). - UI Changes: Reddit updates its interface occasionally—use your browser’s DevTools to inspect elements and tweak selectors if needed.
- Hidden Scores: Some comments hide scores (showing "•")—the
parse_scorefunction skips these by returning 0, but you can modify this logic if you want to handle them differently.
Full Working Example
from selenium.webdriver.common.by import By from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def parse_score(score_text): score_text = score_text.strip().lower() if 'k' in score_text: return int(float(score_text.replace('k', '')) * 1000) elif 'm' in score_text: return int(float(score_text.replace('m', '')) * 1000000) else: try: return int(score_text) except ValueError: return 0 # Initialize Chrome driver (replace with your preferred browser) driver = webdriver.Chrome() driver.get("https://www.reddit.com/user/YourTargetUser/comments/") # Wait for initial comments to load WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, '[class^="comment t1_"]')) ) # Optional: Scroll to load more comments (repeat as needed) for _ in range(3): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") WebDriverWait(driver, 5).until( EC.presence_of_element_located((By.CSS_SELECTOR, '[class^="comment t1_"]')) ) # Fetch all comments and find the top-voted one comments = driver.find_elements(By.CSS_SELECTOR, '[class^="comment t1_"]') max_score = -1 top_comment_element = None for comment in comments: score_elem = comment.find_element(By.CSS_SELECTOR, '[class*="score"]') score = parse_score(score_elem.text) if score > max_score: max_score = score top_comment_element = comment # Output results if top_comment_element: comment_content = top_comment_element.find_element(By.CSS_SELECTOR, '[data-testid="comment"]').text print(f"Top-Voted Comment (Score: {max_score}):\n{comment_content}") else: print("No valid comments found.") # Clean up driver.quit()
内容的提问来源于stack exchange,提问作者Luigi Colella

