You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pythonic方法从文件信息列表提取每日最大文件并构建字典

Pythonic way to get daily largest file by size from a file info list

Problem Statement

I have a list of strings containing file date, name, and size information:

tables = ["20180512, name=file01, size=100", "20180512, name=file02, size=90", "20180513, name=file01, size=70", "20180513, name=file02, size=70", "20180513, name=file03, size=80", "20180514, name=file01, size=100", "20180514, name=file02, size=90"]

I need to build a dictionary where keys are dates, and values are the names of the files with the largest size on that date. The expected result is:

{20180512: 'file01', 20180513: 'file03', 20180514: 'file01'}

I know I can do this with nested loops and extra data structures, but I'm looking for a more efficient, Pythonic approach.


Solution 1: Single-pass dictionary update (Most efficient)

This approach only requires one pass through the list (O(n) time complexity) and uses minimal extra memory. We temporarily store tuples of (size, name) in the dictionary to make size comparisons quick, then convert to the final format with a dictionary comprehension:

tables = ["20180512, name=file01, size=100", "20180512, name=file02, size=90", "20180513, name=file01, size=70", "20180513, name=file02, size=70", "20180513, name=file03, size=80", "20180514, name=file01, size=100", "20180514, name=file02, size=90"]

daily_largest = {}
for entry in tables:
    # Split the entry into individual components
    date_str, name_part, size_part = entry.split(', ')
    date = int(date_str)
    filename = name_part.split('=')[1]
    file_size = int(size_part.split('=')[1])
    
    # Update the dictionary: either add the entry or replace if current file is larger
    if date not in daily_largest or file_size > daily_largest[date][0]:
        daily_largest[date] = (file_size, filename)

# Convert to the desired dictionary format
result = {date: filename for date, (size, filename) in daily_largest.items()}
print(result)

Why this works: We avoid unnecessary loops by keeping track of the largest file we've seen so far for each date as we iterate. Storing the size alongside the name makes comparisons straightforward and fast.


Solution 2: Using itertools.groupby (More declarative)

If you prefer a functional programming style, itertools.groupby is a clean way to group entries by date, then find the maximum size entry in each group. Note that we need to sort the entries first (since groupby only groups consecutive matching keys):

from itertools import groupby

tables = ["20180512, name=file01, size=100", "20180512, name=file02, size=90", "20180513, name=file01, size=70", "20180513, name=file02, size=70", "20180513, name=file03, size=80", "20180514, name=file01, size=100", "20180514, name=file02, size=90"]

# First, process each entry into a tuple of (date, filename, size)
processed_entries = []
for entry in tables:
    date_str, name_part, size_part = entry.split(', ')
    processed_entries.append(
        (int(date_str), 
         name_part.split('=')[1], 
         int(size_part.split('=')[1]))
    )

# Sort entries by date (required for groupby to work correctly)
processed_entries.sort(key=lambda x: x[0])

# Group by date and find the largest file in each group
result = {}
for date, group in groupby(processed_entries, key=lambda x: x[0]):
    # Use max() with a key to get the entry with the largest size
    largest_file = max(group, key=lambda x: x[2])
    result[date] = largest_file[1]

print(result)

Tradeoff: This approach has an O(n log n) time complexity due to the sort step, which makes it slightly less efficient for very large datasets. However, the code reads more like a description of what you're doing ("group by date, then take the max size file") which can be easier to maintain.


Final Notes

Both approaches are Pythonic—choose the first one if raw performance is your priority, and the second if you prefer a more declarative, readable style.

内容的提问来源于stack exchange,提问作者maynull

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:39:59