You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Python爬虫输出的CSV数据按分组合并格式整理?

问题:如何将爬虫导出的CSV数据按分组合并格式整理?

我的爬虫代码:

data = []

while True:
    print(url)
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.content, 'html.parser')
    links = soup.select_one('li.page-item.nb.active')
    
    for links in soup.find_all("h6", {"class": "text-primary title"}):
        sublink = links.find("a").get("href")
        new_link = "LINK" + sublink
        response2 = requests.get(new_link)
        soup2 = BeautifulSoup(response2.content, 'html.parser')
        
        # print('-------------------')
        heading = soup2.find('h1').text
        print(heading)

        table = soup2.find_all('tbody')[0]
        for i in table.find_all('td', class_='title'):
            movies = i.find('a', class_="text-primary")
            for movie in movies:
                data.append((heading,movie))
                
        df = pd.DataFrame(data=data)
        df.to_csv('list.csv', index=False, encoding='utf-8')

    next_page = soup.select_one('li.page-item.next>a')
    if next_page:
        next_url = next_page.get('href')
        url = urljoin(url, next_url)
    else:
        break

当前CSV输出格式:

Column1,Column2
James,movie 1
James,movie 2
James,movie 3

期望输出格式:

Column1,Column2
James,Movie1, Movie2, Movie3
Peter,Movie1, Movie2, Movie3

解决方案:

你可以通过修改数据收集逻辑,用字典按heading分组存储电影,最后统一生成CSV,直接得到目标格式,同时提升爬虫效率。

修改后的代码:

# 改用字典存储,key为heading,value为对应电影列表
data = {}

while True:
    print(url)
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.content, 'html.parser')
    links = soup.select_one('li.page-item.nb.active')
    
    for links in soup.find_all("h6", {"class": "text-primary title"}):
        sublink = links.find("a").get("href")
        new_link = "LINK" + sublink
        response2 = requests.get(new_link)
        soup2 = BeautifulSoup(response2.content, 'html.parser')
        
        heading = soup2.find('h1').text.strip()  # 去除首尾空白,避免分组错误
        print(heading)

        table = soup2.find_all('tbody')[0]
        # 初始化当前heading的电影列表
        if heading not in data:
            data[heading] = []
        
        for i in table.find_all('td', class_='title'):
            movie_tag = i.find('a', class_="text-primary")
            if movie_tag:  # 防止找不到标签报错
                movie_name = movie_tag.text.strip()
                data[heading].append(movie_name)
    
    next_page = soup.select_one('li.page-item.next>a')
    if next_page:
        next_url = next_page.get('href')
        url = urljoin(url, next_url)
    else:
        break

# 将字典转换为DataFrame所需格式
final_data = []
for name, movies in data.items():
    movies_str = ", ".join(movies)  # 把电影列表拼接成字符串
    final_data.append([name, movies_str])

# 生成并导出CSV
df = pd.DataFrame(final_data, columns=['Column1', 'Column2'])
df.to_csv('list.csv', index=False, encoding='utf-8')

关键修改说明:

  1. 存储结构优化:用字典替代列表,自动按heading分组,避免后续额外的合并操作。
  2. 减少IO操作:原代码每次循环都写入CSV,现在改为爬完所有页面后统一导出,提升效率。
  3. 异常防护:添加if movie_tag:判断,避免页面结构异常时因找不到标签崩溃。
  4. 字符串清理:用.strip()去除标题和电影名的首尾空白,避免因空格导致的分组错误。

补充:已有原始CSV的快速处理方法

如果不想修改爬虫代码,也可以直接读取已生成的list.csv,用pandas的分组功能合并数据:

import pandas as pd

# 读取原始CSV
df = pd.read_csv('list.csv')
# 按Column1分组,拼接Column2的内容
df_grouped = df.groupby('Column1')['Column2'].apply(', '.join).reset_index()
# 导出合并后的CSV
df_grouped.to_csv('list_grouped.csv', index=False, encoding='utf-8')

内容的提问来源于stack exchange,提问作者nidiv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 13:50:18