如何为列表元素分配分组序号并创建DataFrame?
Solution to Generate Grouped Sequence DataFrame
Hey there, here are a couple of simple and efficient ways to achieve your desired DataFrame where every two consecutive files share a sequence number, and the last lone file gets its own unique number.
Method 1: Use Index Integer Division (Most Straightforward)
This approach leverages the DataFrame's default 0-based index. By doing integer division of each index by 2 and adding 1, we get exactly the grouping pattern you need.
import pandas as pd # Your original file list file_list = ['file1.txt', 'file2.txt', 'file3.txt', 'file4.txt', 'file5.txt', 'file6.txt', 'file7.txt', 'file8.txt', 'file9.txt', 'file10.txt', 'file11.txt', 'file12.txt', 'file13.txt', 'file14.txt', 'file15.txt'] # Create DataFrame with filenames df = pd.DataFrame({'filename': file_list}) # Generate the sequence column df['sequence'] = (df.index // 2) + 1 # Check the result print(df)
How it works:
- The index starts at 0: for indices 0 & 1,
0//2 = 0→ add 1 gives sequence 1 - Indices 2 & 3 →
2//2 =1→ add 1 gives sequence 2, and so on - The last element (index 14) →
14//2=7→ add 1 gives sequence 8, which matches your requirement perfectly.
Method 2: Using NumPy for Explicit Sequence Generation
If you prefer a more explicit approach to build the sequence list first, NumPy's repeat function makes this easy:
import pandas as pd import numpy as np file_list = ['file1.txt', 'file2.txt', 'file3.txt', 'file4.txt', 'file5.txt', 'file6.txt', 'file7.txt', 'file8.txt', 'file9.txt', 'file10.txt', 'file11.txt', 'file12.txt', 'file13.txt', 'file14.txt', 'file15.txt'] # Calculate total number of sequences: (15 elements → 8 sequences) total_sequences = (len(file_list) + 1) // 2 # Equivalent to ceiling division # Create sequence array where each number repeats twice, then truncate to match list length sequence_array = np.repeat(np.arange(1, total_sequences +1), 2)[:len(file_list)] # Build the DataFrame df = pd.DataFrame({'filename': file_list, 'sequence': sequence_array}) print(df)
How it works:
(len(file_list)+1)//2gives us the total number of sequences (ceiling division ensures we account for the last lone element)np.repeatduplicates each sequence number twice, then we slice the array to match the length of our file list, so the last element only gets the final sequence number once.
Both methods will produce exactly the DataFrame you outlined. The first method is more concise and efficient for most cases!
内容的提问来源于stack exchange,提问作者Shivika Patel
相关产品推荐
相关产品推荐

