You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python/BeautifulSoup爬取Electrive文章末尾多来源链接的技术需求及代码优化问询

Hey there! Let's tweak your code to properly extract and format the article sources as you need, handling 1-3 sources seamlessly. Here's the optimized solution:

from bs4 import BeautifulSoup
import requests

URL = 'https://www.electrive.com/2022/02/20/byd-planning-model-3-like-800-volt-sedan-called-seal/'

# Fetch and parse the page
response = requests.get(URL)
response.raise_for_status()
soup = BeautifulSoup(response.text, 'lxml')

# Locate the "Sources:" label paragraph
sources_label = soup.find('p', string=lambda text: text and 'Sources:' in text.strip())

if sources_label:
    # Get the next paragraph containing source links
    sources_container = sources_label.find_next_sibling('p')
    if sources_container:
        # Extract all source links
        source_links = sources_container.find_all('a')
        # Format each link to match your desired output
        formatted_sources = [f"{link.get_text(strip=True)} ({link['href']})" for link in source_links]
        # Join multiple sources with a Chinese comma
        print('、'.join(formatted_sources))
else:
    print("No sources found for this article.")

Key Improvements:

  • Robust Source Targeting: Instead of slicing arbitrary <a> tags (which breaks if the page structure changes), we directly find the "Sources:" label and its adjacent paragraph holding the links. This works reliably even if the site adds other links in the article body.
  • Dynamic Source Count Support: The code automatically collects all sources in the container, whether there are 1, 2, or 3. No hardcoded indexes needed!
  • Exact Formatting: Each source is formatted to match your requirement: [source text] ([URL]), and multiple sources are joined with a Chinese顿号 for consistency with your example.
  • Error Resilience: Added checks to handle cases where the sources section might be missing, preventing unexpected crashes.

When you run this code for your target URL, you'll get the output:

cnevpost.com (https://cnevpost.com/2022/02/18/byd-seal-set-to-become-new-tesla-model-3-challenger/)

For articles with multiple sources, it will output something like:

xchuxing.com (https://xchuxing.com/article/45850)、cnevpost.com (https://cnevpost.com/2022/02/18/byd-seal-set-to-become-new-tesla-model-3-challenger/)

内容的提问来源于stack exchange,提问作者webscrapeartist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 14:47:39