使用Python/BeautifulSoup爬取Electrive文章末尾多来源链接的技术需求及代码优化问询
Hey there! Let's tweak your code to properly extract and format the article sources as you need, handling 1-3 sources seamlessly. Here's the optimized solution:
from bs4 import BeautifulSoup import requests URL = 'https://www.electrive.com/2022/02/20/byd-planning-model-3-like-800-volt-sedan-called-seal/' # Fetch and parse the page response = requests.get(URL) response.raise_for_status() soup = BeautifulSoup(response.text, 'lxml') # Locate the "Sources:" label paragraph sources_label = soup.find('p', string=lambda text: text and 'Sources:' in text.strip()) if sources_label: # Get the next paragraph containing source links sources_container = sources_label.find_next_sibling('p') if sources_container: # Extract all source links source_links = sources_container.find_all('a') # Format each link to match your desired output formatted_sources = [f"{link.get_text(strip=True)} ({link['href']})" for link in source_links] # Join multiple sources with a Chinese comma print('、'.join(formatted_sources)) else: print("No sources found for this article.")
Key Improvements:
- Robust Source Targeting: Instead of slicing arbitrary
<a>tags (which breaks if the page structure changes), we directly find the "Sources:" label and its adjacent paragraph holding the links. This works reliably even if the site adds other links in the article body. - Dynamic Source Count Support: The code automatically collects all sources in the container, whether there are 1, 2, or 3. No hardcoded indexes needed!
- Exact Formatting: Each source is formatted to match your requirement:
[source text] ([URL]), and multiple sources are joined with a Chinese顿号 for consistency with your example. - Error Resilience: Added checks to handle cases where the sources section might be missing, preventing unexpected crashes.
When you run this code for your target URL, you'll get the output:
cnevpost.com (https://cnevpost.com/2022/02/18/byd-seal-set-to-become-new-tesla-model-3-challenger/)
For articles with multiple sources, it will output something like:
xchuxing.com (https://xchuxing.com/article/45850)、cnevpost.com (https://cnevpost.com/2022/02/18/byd-seal-set-to-become-new-tesla-model-3-challenger/)
内容的提问来源于stack exchange,提问作者webscrapeartist
相关产品推荐
相关产品推荐

