按Type of region列值分组行至列表并存入容器的技术问询
Got it, let's break down how to tackle this problem smoothly—especially with 700k rows, we need solutions that are both fast and memory-friendly. I'll cover Python (with and without Pandas) and C++ since you mentioned containers like vectors/lists.
Option 1: Python with Pandas (Best for Large Datasets)
Pandas is optimized for tabular data and handles 700k rows effortlessly. Here's how to group and store your data:
Steps:
- Read your data into a DataFrame, making sure to handle the correct delimiter (looks like whitespace/tabs in your example).
- Group the DataFrame by the
Type of regioncolumn. - Convert each group into a list of rows, then store these groups in your preferred container (dict or list).
import pandas as pd # Load the data—adjust sep if your file uses tabs instead of spaces df = pd.read_csv("your_data_file.txt", sep="\s+", header=0) # Group into a dictionary: key = region type, value = list of rows (as dicts) grouped_data = df.groupby("Type of region").apply(lambda x: x.to_dict("records")).to_dict() # If you want a list container (each element is a group's row list) grouped_list = [group.to_dict("records") for _, group in df.groupby("Type of region")]
Why this works: Pandas' groupby is highly optimized, so it’ll process your 700k rows in seconds. Converting rows to dictionaries makes each row easy to work with later.
Option 2: Pure Python (No External Dependencies)
If you prefer avoiding Pandas, use the standard library’s csv module and collections.defaultdict to handle grouping:
import csv from collections import defaultdict # Initialize a dictionary where each key maps to a list of rows region_groups = defaultdict(list) # Open and read the file—use delimiter="\t" if your data is tab-separated with open("your_data_file.txt", "r") as f: # Use DictReader to access columns by name reader = csv.DictReader(f, delimiter=" ") for row in reader: region_type = row["Type of region"] region_groups[region_type].append(row) # Store groups in a list container (each element is a list of rows for a region type) grouped_list = list(region_groups.values())
Note: Double-check the delimiter—if your data uses tabs instead of spaces, swap delimiter=" " for delimiter="\t".
Option 3: C++ (Using Vectors)
If you’re working in C++ and want to use vector containers, here’s a streamlined approach:
First, define a struct to hold each row’s data, then use an unordered_map to group rows by region type, and finally collect those groups into a vector of vectors.
#include <iostream> #include <fstream> #include <vector> #include <unordered_map> #include <sstream> #include <string> // Struct to store each row's data—match the columns from your dataset struct RowData { std::string chr; long long cs; long long ce; std::string clone_name; int score; char strand; int locs_per_clone; int capreg_alignments; std::string region_type; }; int main() { // Map region types to their respective row vectors std::unordered_map<std::string, std::vector<RowData>> region_groups; std::ifstream data_file("your_data_file.txt"); std::string line; // Skip the header row std::getline(data_file, line); // Parse each row while (std::getline(data_file, line)) { std::istringstream line_stream(line); RowData row; // Parse values in the order of your columns line_stream >> row.chr >> row.cs >> row.ce >> row.clone_name >> row.score >> row.strand >> row.locs_per_clone >> row.capreg_alignments >> row.region_type; // Add the row to the corresponding group region_groups[row.region_type].push_back(row); } // Collect all groups into a single vector of vectors std::vector<std::vector<RowData>> grouped_vectors; for (auto& [region, rows] : region_groups) { grouped_vectors.push_back(std::move(rows)); } // Your grouped data is now in grouped_vectors—use as needed return 0; }
Key Notes:
- Use
long longfor numeric columns likecsandceto avoid integer overflow. - The
std::movecall optimizes memory usage by transferring ownership of the row vectors instead of copying them.
内容的提问来源于stack exchange,提问作者Yujin Kim

