如何在Perl中删除表格里仅包含1的列?
需求:删除表格中全为1的列
原始表格如下:
| head | v1 | v2 | v3 | v4 | v5 | v6 |
|---|---|---|---|---|---|---|
| stn2 | 1 | 4 | 1 | 1 | 4 | 2 |
| stn2 | 1 | 4 | 1 | 1 | 4 | 2 |
| stn3 | 1 | 4 | 1 | 1 | 4 | 2 |
| stn4 | 1 | 4 | 1 | 1 | 4 | 3 |
| stn4 | 1 | 4 | 1 | 1 | 4 | 2 |
| stn5 | 1 | 4 | 1 | 1 | 4 | 4 |
| stn6 | 1 | 3 | 1 | 1 | 4 | 3 |
| stn7 | 4 | 4 | 1 | 1 | 4 | 4 |
| stn8 | 4 | 4 | 1 | 1 | 4 | 3 |
| stn9 | 2 | 4 | 1 | 1 | 4 | 3 |
我想要删除所有仅包含1的列,目前用了一段冗长的代码把六列分别存入数组:
#!/usr/bin/perl use strict; use warnings; use Data::Dumper; use feature 'say'; open(my $RSKF, " < ../risks.txt") || die "open risks.txt: failed $! ($^E)"; my $line; my $count; my @one_column; my @two_column; my @three_column; my @four_column; my @five_column; my @six_column; my $one_column; my $two_column; my $three_column; my $four_column; my $five_column; my $six_column; my @remove; my @keep1; my $keepcount=0; while($line = <$RSKF>){ push(@one_column, (split(/\s+/, $line))[2]); push(@two_column, (split(/\s+/, $line))[3]); push(@three_column, (split(/\s+/, $line))[4]); push(@four_column, (split(/\s+/, $line))[5]); push(@five_column, (split(/\s+/, $line))[6]); push(@six_column, (split(/\s+/, $line))[7]); }
我尝试遍历第四列的代码无法正常工作:
$count=0; for (my $i=1; $i < @four_column; ++$i){ if($four_column[$i] ge '2'){ $count++; } if($count > 0){ @four_column=@keep1; $keepcount++; } else{ @four_column=@remove; }
肯定有更简便的实现方法,恳请帮助。
简洁实现方案
不需要单独为每列创建数组,我们可以先读取所有行数据,标记出非全1的列,最后只输出这些列即可。
完整代码
#!/usr/bin/perl use strict; use warnings; # 打开文件 open(my $fh, '<', '../risks.txt') or die "无法打开文件: $!"; # 读取所有行,转换为二维数组 my @rows; while (my $line = <$fh>) { chomp $line; push @rows, [ split /\s+/, $line ]; } close $fh; # 筛选需要保留的列(排除全为1的列) my @keep_cols; my $total_cols = scalar @{$rows[0]}; for my $col_idx (0 .. $total_cols - 1) { my $is_all_one = 1; # 检查当前列的所有行 for my $row (@rows) { if ($row->[$col_idx] ne '1') { $is_all_one = 0; last; } } # 不是全1的列加入保留列表 push @keep_cols, $col_idx unless $is_all_one; } # 输出结果 for my $row (@rows) { print join("\t", @{$row}[@keep_cols]) . "\n"; }
代码说明
- 数据存储:将整个表格转为二维数组
@rows,每一行是一个数组引用,操作更灵活,避免了单独定义多列数组的冗余。 - 列筛选:遍历每一列,只要该列存在非1的元素,就标记为需要保留。
- 结果输出:遍历所有行,只输出保留的列,用制表符分隔维持表格格式。
原代码问题说明
- 单独为每列创建数组的方式扩展性差,列数变化时需要修改大量代码。
- 遍历第四列的逻辑错误:循环内部重复执行列保留/删除操作,且
@keep1是空数组,赋值后会清空目标列数据。 - 变量名重复(如同时存在数组
@four_column和标量$four_column),容易引发逻辑混淆。
内容的提问来源于stack exchange,提问作者Zilore Mumba
相关产品推荐
相关产品推荐

