You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页新闻爬取求助:单篇文章段落合并失败及列表存储异常

问题分析与解决方案

我来帮你搞定这个问题!你的核心问题出在**paragraphtext列表的初始化时机**上——这个列表在函数开头就定义了,属于整个scrape函数的全局变量,每次处理新文章时你只是往里面追加段落,却没有清空它,导致所有文章的段落都累积到同一个列表里。最终thearticle里的每个元素都是所有文章段落的总和,所以myarticle[0]自然会输出所有文章的内容。

修改后的完整代码

import requests
from bs4 import BeautifulSoup as bs
from time import time, sleep
from random import randint
from warnings import warn
from IPython.display import clear_output

def scrape(url):
    user_agent = {'user-agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64; Trident/7.0; Touch; rv:11.0) like Gecko'}
    request = 0
    urls = [f"{url}{x}" for x in range(1,2)]
    params = { "orderby": "relevance", }
    pagelinks = []
    title = []
    thearticle = []
    
    for page in urls:
        response = requests.get(url=page, headers=user_agent, params=params)
        # controlling the crawl-rate
        start_time = time()
        #pause the loop
        sleep(randint(8,15))
        #monitor the requests
        request += 1
        elapsed_time = time() - start_time
        print('Request:{}; Frequency: {} request/s'.format(request, request/elapsed_time))
        clear_output(wait = True)
        #throw a warning for non-200 status codes
        if response.status_code != 200:
            warn('Request: {}; Status code: {}'.format(request, response.status_code))
        #Break the loop if the number of requests is greater than expected
        if request > 72:
            warn('Number of request was greater than expected.')
            break
        #parse the content
        soup_page = bs(response.text, 'lxml')
        #select all the articles for a single page
        containers = soup_page.findAll("li", {'class': 'article'})
        #scrape the links of the articles
        for i in containers:
            url = i.find('a')
            pagelinks.append(url.get('href'))
        #scrape the titles of the articles
        for i in containers:
            atitle = i.find(class_ = 'entry-heading').find('a')
            thetitle = atitle.get_text()
            title.append(thetitle)
    
    for pagelink in pagelinks:
        # 关键修改:每处理一篇新文章,重新初始化段落列表
        paragraphtext = []
        #get page text
        page = requests.get(pagelink)
        #parse with BeautifulSoup
        soup = bs(page.text, 'lxml')
        containerr = soup.find("div", class_=['entry-content', 'entry-content-read-more'])
        articletext = containerr.find_all('p')
        for paragraph in articletext:
            #get the text only
            text = paragraph.get_text()
            paragraphtext.append(text)
        #combine all paragraphs into an article
        thearticle.append(paragraphtext)
    
    # join paragraphs to re-create the article
    myarticle = [''.join(article) for article in thearticle]
    print(myarticle[0])
    return myarticle  # 建议返回结果,方便后续使用

print(scrape('https://nypost.com/search/China+COVID-19/page/'))

核心修改点说明

  • 将paragraphtext = []移动到for pagelink in pagelinks:循环的内部,这样每处理一篇新文章时,都会创建一个全新的空列表来存储当前文章的段落,彻底避免了段落的跨文章累积。
  • 额外加了return myarticle的建议,这样你调用函数后可以直接拿到整理好的文章列表,方便后续存储或分析。

现在运行修改后的代码,myarticle[0]就只会输出第一篇文章的完整内容了,每一篇文章都会被正确独立存储到列表中。

内容的提问来源于stack exchange,提问作者Yue Peng

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 18:37:46