You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用awk等工具处理30GB大文件生成对称OD矩阵?

大OD矩阵对称化处理方案

现有一个30GB的OD(起点-终点)矩阵,以列表形式存储在inputfile.csv中,示例输入如下:

"origin_id","destination_id","trips"
"0","0","20"
"0","1","12"
"0","2","8"
"1","0","23"
"1","1","50"
"1","2","6"
"2","1","9"
"2","2","33"

该文件仅记录出行量非零的OD对,出行量为0的OD对未被保存。需求是计算对称矩阵S=(OD+DO)/2:即每个OD对的出行量取自身与反向DO对的平均值;若反向DO对不存在,则取自身的一半。期望输出如下:

"origin_id","destination_id","trips"
"0","0","20"  
"0","1","17.5"
"0","2","4"
"1","1","50"
"1","2","7.5"
"2","2","33"

AWK实现方案

针对大文件优化,先将所有数据存入关联数组,再遍历计算对称值,避免重复输出:

BEGIN {
    FS = ","
    OFS = ","
    print "\"origin_id\",\"destination_id\",\"trips\""
}

# 读取所有行,存入数组,同时记录所有出现过的起点/终点
NR > 1 {
    gsub(/"/, "", $1); gsub(/"/, "", $2); gsub(/"/, "", $3)
    key = $1 "," $2
    data[key] = $3
    origins[$1] = 1
    destinations[$2] = 1
}

# 遍历所有OD对,仅处理起点<=终点的情况,避免重复计算
END {
    for (o in origins) {
        for (d in destinations) {
            if (o > d) continue
            key = o "," d
            rev_key = d "," o
            val = 0
            if (key in data) val += data[key]
            if (rev_key in data) val += data[rev_key]
            val = val / 2
            if (val > 0) {
                print "\""o"\",\""d"\",\""val"\""
            }
        }
    }
}

运行命令:

awk -f symmetrize_od.awk inputfile.csv > outputfile.csv

注:若起点/终点范围极大(如百万级),双重循环会较慢,但AWK对大文件读取效率极高,内存占用仅取决于非零OD对的数量。


Python实现方案(内存友好型)

通过字典存储所有OD对,遍历一次字典即可完成计算,避免冗余循环:

import csv

def symmetrize_od(input_path, output_path):
    od_data = {}
    # 读取所有非零OD对到字典
    with open(input_path, 'r', encoding='utf-8') as infile:
        reader = csv.DictReader(infile)
        for row in reader:
            o = row['origin_id'].strip('"')
            d = row['destination_id'].strip('"')
            trips = float(row['trips'].strip('"'))
            od_data[(o, d)] = trips
    
    processed = set()
    # 写入对称化后的结果
    with open(output_path, 'w', encoding='utf-8', newline='') as outfile:
        writer = csv.writer(outfile, quoting=csv.QUOTE_ALL)
        writer.writerow(['origin_id', 'destination_id', 'trips'])
        for (o, d), trips12 in od_data.items():
            if (o, d) in processed:
                continue
            # 计算对称值
            if (d, o) in od_data:
                trips21 = od_data[(d, o)]
                sym_val = (trips12 + trips21) / 2
                writer.writerow([o, d, sym_val])
                if o != d:
                    processed.add((d, o))
            else:
                sym_val = trips12 / 2
                writer.writerow([o, d, sym_val])
            processed.add((o, d))

if __name__ == '__main__':
    symmetrize_od('inputfile.csv', 'outputfile.csv')

注:该方案内存占用仅为非零OD对的数量,适合千万级以内的非零对场景,若非零对数量超出内存,可考虑分块处理。


内容的提问来源于stack exchange,提问作者ElTitoFranki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 04:02:47