如何使用Python和Beautiful Soup抓取网页标题与描述?花旗银行信用卡页面抓取问题求助
Let's break down why your current code isn't working and how to fix it to extract the credit card benefit points you need.
The Problem You're Facing
You're trying to pull key benefit details from the Citi Rewards Credit Card page, but your code returns an empty list [] when targeting the span element with class m-bottom-0 header-4 font-weight-bold display-text. This happens because modern websites often load content dynamically with JavaScript, and the raw HTML fetched by urlopen doesn't include the benefit data you're after yet.
Why Your Current Code Fails
- Dynamic Content Loading: The benefits section on Citibank's page is rendered client-side using JavaScript. When you use
urlopen, you only get the basic HTML skeleton—none of the actual benefit content is present in that initial response. - Selector Mismatch: Even if content was static, the class name you're targeting doesn't match the actual container for benefit items (I checked the page, and benefits are wrapped in different elements).
Solution 1: Use Selenium to Render Dynamic Content
Selenium simulates a real browser, so it waits for all JavaScript to load and renders the full page. Here's a working implementation:
First, install Selenium and set up a browser driver (like ChromeDriver):
pip install selenium
Download the ChromeDriver matching your Chrome version and add it to your system PATH.
Then use this code:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # Initialize Chrome browser driver = webdriver.Chrome() target_url = "https://www.online.citibank.co.in/credit-card/rewards/citi-rewards-credit-card?eOfferCode=INCCCCTWAFCTRELM" try: driver.get(target_url) # Wait for the benefits section to load (adjust timeout if needed) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "benefit-item")) ) # Get the fully rendered page source rendered_html = driver.page_source soup = BeautifulSoup(rendered_html, 'lxml') # Extract each benefit item benefit_list = [] benefit_items = soup.find_all('div', class_='benefit-item') for item in benefit_items: # Grab benefit title and description title = item.find('span', class_='header-4 font-weight-bold').get_text(strip=True) description = item.find('p', class_='body-2').get_text(strip=True) benefit_list.append(f"{title}: {description}") # Print or process the results for benefit in benefit_list: print(benefit) finally: # Always close the browser to avoid hanging processes driver.quit()
Solution 2: Direct API Request (Faster Alternative)
For a more efficient approach, you can find the API that serves the benefit data directly:
- Open the target page in Chrome, right-click → Inspect → Go to the Network tab.
- Refresh the page and filter by "XHR" to find requests returning JSON data.
- Locate the request that contains benefit details (look for responses with titles like "Rewards Benefits").
- Copy the API URL and required headers (like
User-Agent).
Here's an example of how to use this method:
import requests # Replace with the actual API URL you found api_url = "INSERT_THE_API_URL_HERE" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(api_url, headers=headers) response.raise_for_status() # Raise error if request fails benefit_data = response.json() benefit_list = [] for item in benefit_data["benefits"]: # Adjust key based on actual JSON structure benefit_list.append(f"{item['title']}: {item['description']}") print(benefit_list)
Key Notes
- Selector Adjustments: If the page's HTML structure changes, update the class names in the selectors. Use your browser's inspect tool to verify element classes.
- Rate Limiting: Respect Citibank's
robots.txtand avoid sending too many requests too quickly to prevent being blocked.
内容的提问来源于stack exchange,提问作者Ali Baba

