You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Python爬虫输出的多个列表合并为单列并统计各值出现次数

实现方案
  • 原代码核心问题是循环过程中没有留存所有提取到的背景值,每次遍历都覆盖了traits变量,同时直接打印单次结果,自然输出的是多组独立列表
  • 只需额外新增一个全局的汇总列表,每次提取到背景值就追加到这个列表里,后续统一处理即可

完整可运行代码

from bs4 import BeautifulSoup
import requests
import pandas as pd
import numpy as np

# 初始化汇总背景值的空列表
all_background = []
page = 1

while page != 10:
    content = requests.get('https://raw.githubusercontent.com/recklesslabs/wickedcraniums/main/{}'.format(page))
    soup = BeautifulSoup(content.text, 'html.parser')
    page = page + 1
    
    dic_list = list(map(eval, soup))
    for dic in dic_list:
        traits = dic["attributes"]
    df = pd.DataFrame.from_dict(traits, orient='columns').to_numpy()
    df1 = list(np.concatenate(df[0:1]))
    # 提取背景值追加到汇总列表
    all_background.append(df1[1])

# 输出单列结构化DataFrame
df_background = pd.DataFrame(all_background, columns=['Background'])
print(df_background)

# 统计各背景值出现次数
count_result = df_background['Background'].value_counts().reset_index()
count_result.columns = ['Background', '出现次数']
# 按要求格式打印统计结果
print("\nBackground")
for _, row in count_result.iterrows():
    print(f"{row['Background']} - {row['出现次数']}")

可选原生统计方案

如果不想依赖pandas做统计,也可以用Python原生库实现统计功能:

from collections import Counter
count_dict = Counter(all_background)
print("\nBackground")
for k, v in count_dict.items():
    print(f"{k} - {v}")

内容的提问来源于stack exchange,提问作者intermarketics

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 20:39:05