网页新闻爬取求助:单篇文章段落合并失败及列表存储异常
问题分析与解决方案
我来帮你搞定这个问题!你的核心问题出在**paragraphtext列表的初始化时机**上——这个列表在函数开头就定义了,属于整个scrape函数的全局变量,每次处理新文章时你只是往里面追加段落,却没有清空它,导致所有文章的段落都累积到同一个列表里。最终thearticle里的每个元素都是所有文章段落的总和,所以myarticle[0]自然会输出所有文章的内容。
修改后的完整代码
import requests from bs4 import BeautifulSoup as bs from time import time, sleep from random import randint from warnings import warn from IPython.display import clear_output def scrape(url): user_agent = {'user-agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64; Trident/7.0; Touch; rv:11.0) like Gecko'} request = 0 urls = [f"{url}{x}" for x in range(1,2)] params = { "orderby": "relevance", } pagelinks = [] title = [] thearticle = [] for page in urls: response = requests.get(url=page, headers=user_agent, params=params) # controlling the crawl-rate start_time = time() #pause the loop sleep(randint(8,15)) #monitor the requests request += 1 elapsed_time = time() - start_time print('Request:{}; Frequency: {} request/s'.format(request, request/elapsed_time)) clear_output(wait = True) #throw a warning for non-200 status codes if response.status_code != 200: warn('Request: {}; Status code: {}'.format(request, response.status_code)) #Break the loop if the number of requests is greater than expected if request > 72: warn('Number of request was greater than expected.') break #parse the content soup_page = bs(response.text, 'lxml') #select all the articles for a single page containers = soup_page.findAll("li", {'class': 'article'}) #scrape the links of the articles for i in containers: url = i.find('a') pagelinks.append(url.get('href')) #scrape the titles of the articles for i in containers: atitle = i.find(class_ = 'entry-heading').find('a') thetitle = atitle.get_text() title.append(thetitle) for pagelink in pagelinks: # 关键修改:每处理一篇新文章,重新初始化段落列表 paragraphtext = [] #get page text page = requests.get(pagelink) #parse with BeautifulSoup soup = bs(page.text, 'lxml') containerr = soup.find("div", class_=['entry-content', 'entry-content-read-more']) articletext = containerr.find_all('p') for paragraph in articletext: #get the text only text = paragraph.get_text() paragraphtext.append(text) #combine all paragraphs into an article thearticle.append(paragraphtext) # join paragraphs to re-create the article myarticle = [''.join(article) for article in thearticle] print(myarticle[0]) return myarticle # 建议返回结果,方便后续使用 print(scrape('https://nypost.com/search/China+COVID-19/page/'))
核心修改点说明
- 将
paragraphtext = []移动到for pagelink in pagelinks:循环的内部,这样每处理一篇新文章时,都会创建一个全新的空列表来存储当前文章的段落,彻底避免了段落的跨文章累积。 - 额外加了
return myarticle的建议,这样你调用函数后可以直接拿到整理好的文章列表,方便后续存储或分析。
现在运行修改后的代码,myarticle[0]就只会输出第一篇文章的完整内容了,每一篇文章都会被正确独立存储到列表中。
内容的提问来源于stack exchange,提问作者Yue Peng
相关产品推荐
相关产品推荐

