如何用BeautifulSoup提取ul下li标签中的福布斯热门文章URL
Hey there! Let's break down why you might be struggling to grab those Forbes hot article URLs and fix it step by step.
1. Check if the content is dynamically loaded
Modern sites like Forbes often use JavaScript to render content on the fly. If you're just using requests.get() to fetch the page, you're probably only getting the initial empty HTML shell—those <li> tags you need haven't been loaded yet.
Solution: Use tools that can execute JavaScript
Try selenium or requests-html to get the fully rendered page. Here's a quick selenium example:
from selenium import webdriver from bs4 import BeautifulSoup import time # Initialize Chrome driver (make sure you have the matching chromedriver installed) driver = webdriver.Chrome() driver.get("https://www.forbes.com/") # Wait a few seconds for the page to load (you can use explicit waits for better reliability) time.sleep(3) # Grab the rendered page source page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # Clean up driver.quit()
2. Make sure your selectors are correct
It's easy to target the wrong <ul> tag—maybe there are multiple uls on the page, or you got the class/id wrong.
Solution: Copy the exact selector from dev tools
- Open Forbes in your browser, hit F12 to open DevTools
- Find one of the target
<li>tags, right-click its parent<ul> - Choose
Copy > Copy selectororCopy > Copy XPath - Use that exact selector in your code
Example using a copied CSS selector:
# Locate the target ul tag target_ul = soup.select_one('.popular-articles__list') # Replace with your copied selector if target_ul: article_urls = [] for li in target_ul.find_all('li'): a_tag = li.find('a') if a_tag and 'href' in a_tag.attrs: raw_url = a_tag['href'] # Fix relative URLs if not raw_url.startswith('http'): full_url = f"https://www.forbes.com{raw_url}" else: full_url = raw_url article_urls.append(full_url) print("Got article URLs:", article_urls) else: print("Couldn't find the target ul tag—double-check your selector!")
3. Bypass Forbes' anti-scraping measures
Forbes might block your request if it detects you're a bot. You need to mimic a real browser.
Solution: Add proper request headers
If you're sticking with requests (and the content is static), add a user-agent header:
import requests from bs4 import BeautifulSoup headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } response = requests.get("https://www.forbes.com/", headers=headers) response.raise_for_status() # Catch HTTP errors like 403 soup = BeautifulSoup(response.text, 'html.parser') # Proceed with selector logic as above
4. Verify the page structure hasn't changed
Websites update their HTML structure all the time. Double-check in DevTools that the <li> tags are still nested under the same <ul>—maybe the class name got updated, or the hierarchy shifted.
内容的提问来源于stack exchange,提问作者Daniel Segal

