You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按Type of region列值分组行至列表并存入容器的技术问询

Efficiently Grouping 700k Rows by "Type of Region"

Got it, let's break down how to tackle this problem smoothly—especially with 700k rows, we need solutions that are both fast and memory-friendly. I'll cover Python (with and without Pandas) and C++ since you mentioned containers like vectors/lists.

Option 1: Python with Pandas (Best for Large Datasets)

Pandas is optimized for tabular data and handles 700k rows effortlessly. Here's how to group and store your data:

Steps:

  • Read your data into a DataFrame, making sure to handle the correct delimiter (looks like whitespace/tabs in your example).
  • Group the DataFrame by the Type of region column.
  • Convert each group into a list of rows, then store these groups in your preferred container (dict or list).
import pandas as pd

# Load the data—adjust sep if your file uses tabs instead of spaces
df = pd.read_csv("your_data_file.txt", sep="\s+", header=0)

# Group into a dictionary: key = region type, value = list of rows (as dicts)
grouped_data = df.groupby("Type of region").apply(lambda x: x.to_dict("records")).to_dict()

# If you want a list container (each element is a group's row list)
grouped_list = [group.to_dict("records") for _, group in df.groupby("Type of region")]

Why this works: Pandas' groupby is highly optimized, so it’ll process your 700k rows in seconds. Converting rows to dictionaries makes each row easy to work with later.

Option 2: Pure Python (No External Dependencies)

If you prefer avoiding Pandas, use the standard library’s csv module and collections.defaultdict to handle grouping:

import csv
from collections import defaultdict

# Initialize a dictionary where each key maps to a list of rows
region_groups = defaultdict(list)

# Open and read the file—use delimiter="\t" if your data is tab-separated
with open("your_data_file.txt", "r") as f:
    # Use DictReader to access columns by name
    reader = csv.DictReader(f, delimiter=" ")
    for row in reader:
        region_type = row["Type of region"]
        region_groups[region_type].append(row)

# Store groups in a list container (each element is a list of rows for a region type)
grouped_list = list(region_groups.values())

Note: Double-check the delimiter—if your data uses tabs instead of spaces, swap delimiter=" " for delimiter="\t".

Option 3: C++ (Using Vectors)

If you’re working in C++ and want to use vector containers, here’s a streamlined approach:

First, define a struct to hold each row’s data, then use an unordered_map to group rows by region type, and finally collect those groups into a vector of vectors.

#include <iostream>
#include <fstream>
#include <vector>
#include <unordered_map>
#include <sstream>
#include <string>

// Struct to store each row's data—match the columns from your dataset
struct RowData {
    std::string chr;
    long long cs;
    long long ce;
    std::string clone_name;
    int score;
    char strand;
    int locs_per_clone;
    int capreg_alignments;
    std::string region_type;
};

int main() {
    // Map region types to their respective row vectors
    std::unordered_map<std::string, std::vector<RowData>> region_groups;
    std::ifstream data_file("your_data_file.txt");
    std::string line;

    // Skip the header row
    std::getline(data_file, line);

    // Parse each row
    while (std::getline(data_file, line)) {
        std::istringstream line_stream(line);
        RowData row;
        
        // Parse values in the order of your columns
        line_stream >> row.chr >> row.cs >> row.ce >> row.clone_name >> row.score >> row.strand
                    >> row.locs_per_clone >> row.capreg_alignments >> row.region_type;
        
        // Add the row to the corresponding group
        region_groups[row.region_type].push_back(row);
    }

    // Collect all groups into a single vector of vectors
    std::vector<std::vector<RowData>> grouped_vectors;
    for (auto& [region, rows] : region_groups) {
        grouped_vectors.push_back(std::move(rows));
    }

    // Your grouped data is now in grouped_vectors—use as needed
    return 0;
}

Key Notes:

  • Use long long for numeric columns like cs and ce to avoid integer overflow.
  • The std::move call optimizes memory usage by transferring ownership of the row vectors instead of copying them.

内容的提问来源于stack exchange,提问作者Yujin Kim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:11:25