如何使用BeautifulSoup获取指定位置的特定p标签内容?
<p> Tag Content with BeautifulSoup Hey there! I see the issue—your current find('p') call only grabs the first matching <p> tag, which is why you're getting "2021" instead of the date you want. Let's fix that with a couple straightforward approaches:
Approach 1: Use find_all() and Indexing
Since find_all() returns a list of all matching elements, you can just grab the second item (remember, Python uses 0-based indexing, so index 1 is the second element). We'll also narrow it down to only <p> tags with the bodytext class to avoid unexpected matches:
import nltk import requests from bs4 import BeautifulSoup import time from datetime import date, datetime, timedelta url = 'https://www.richter-helm.eu/index.php?id=26' page = requests.get(url) soup = BeautifulSoup(page.content,'html.parser') results_date = soup.find_all(class_='csc-textpic-text')[1] # Get all p.bodytext tags, then pick the second one p_tags = results_date.find_all('p', class_='bodytext') if len(p_tags) >= 2: # Guard against index errors if the structure changes target_date = p_tags[1].get_text(strip=True) print(target_date) # Output: 10. March 2021 else: print("Couldn't locate the second p tag with class 'bodytext'")
Approach 2: Use CSS Selectors with select_one()
CSS selectors let you target elements directly by their position. The :nth-of-type(2) selector picks the second <p> tag inside your results_date container:
# ... (your existing code up to defining results_date) target_date = results_date.select_one('p.bodytext:nth-of-type(2)').get_text(strip=True) print(target_date) # Output: 10. March 2021
Both methods will reliably pull the "10. March 2021" text you need. The strip=True in get_text() cleans up any extra whitespace around the date for a cleaner output.
内容的提问来源于stack exchange,提问作者maxpython

