基于Python与BS4:如何通过h2类获取对应a标签链接及文本
Solution for Your New Scraping Task
Got it, let's tweak your code to work with the new HTML structure you're dealing with. Since the <h2> has a clear class (top-sec-title) and it's nested inside the <a> tag you need the href from, we can approach this by first targeting the <h2> elements, then grabbing their parent <a> tags.
Here's the adjusted code:
from bs4 import BeautifulSoup # Assume your HTML content is stored in a variable called html_content soup = BeautifulSoup(html_content, 'html.parser') # Find all h2 elements with the target class news_headings = soup.find_all('h2', class_='top-sec-title') for heading in news_headings: # Get the parent <a> tag of the h2 parent_link = heading.parent # Extract the href attribute from the a tag article_url = parent_link.get('href') # Extract the cleaned text from the h2 article_title = heading.get_text(strip=True) print(f"Article URL: {article_url}") print(f"Article Title: {article_title}") print("---")
Key Notes:
- We use
class_='top-sec-title'instead ofattrs={'class': ...}here (both work, butclass_is a cleaner shorthand in BeautifulSoup for targeting class attributes, sinceclassis a reserved keyword in Python). heading.parentdirectly gives us the immediate parent element of the<h2>, which is the<a>tag we need.get_text(strip=True)removes any extra whitespace (like line breaks or leading/trailing spaces) from the heading text, making it cleaner than just.text.
This approach is more reliable for your new HTML structure because the <a> tag doesn't have a unique class to target directly—instead, the <h2> does, so starting there makes sense.
内容的提问来源于stack exchange,提问作者Tayyab Nasir
相关产品推荐
相关产品推荐

