如何提取含pe=1的完整数据条目并生成指定格式表格?
提取含pe=1的完整条目并生成指定表格
需求说明
我有如下格式的文本数据:
>ab:xy_a0by98-2 \Movie= top gun \actor= Tom \Genere=Action \Length=234 \Credits=30 \pe=1 \summry=(Tom|action|234) Top Gun is a 1986 American action drama film directed by Tony Scott, and produced by Don Simpson and Jerry Bruckheimer >ab:xy_b0ha81-5 \Movie= Thor \actor= chris hemsworth \Genere=Action \Length=321 \Credits=20 \pe=0 \summry=(chris|Action|321) Thor embarks on a journey unlike anything he's ever faced a quest for inner peace >ab:xy_c0ma65-1 \Movie= Batman \actor= Bale \Genere=Action \Length=251 \Credits=30 \pe=1 \summry=(Bale|Action|251) From American Psycho to Batman Begins to Vice, Christian Bale is a bonafide A-list star But he missed out on plenty of huge roles along the way. >ab:xy_d0fc78-2 \Movie= Joker \actor= Phoenix \Genere=thriller \Length=341 \Credits=35 \pe=2 \summry=(phoenix|thriller|341) Joker is a 2019 American psychological thriller film directed and produced by Todd Phillips who co-wrote the screenplay with Scott Silver >ab:xy_e0ra81-2 \Movie= Superman \actor= henry cavill \Genere=Action \Length=254 \Credits=28 \pe=1 \summry=(cavill|action|254) Henry William Dalgliesh Cavill is a British actor He is known for his portrayal of Charles Brandon in Showtime's The Tudors
需要完成两个任务:
- 提取所有包含
pe=1的完整条目(每个条目以>开头,包含后续描述内容直到下一个空行),期望输出:
>ab:xy_a0by98-2 \Movie= top gun \actor= Tom \Genere=Action \Length=234 \Credits=30 \pe=1 \summry=(Tom|action|234) Top Gun is a 1986 American action drama film directed by Tony Scott, and produced by Don Simpson and Jerry Bruckheimer >ab:xy_c0ma65-1 \Movie= Batman \actor= Bale \Genere=Action \Length=251 \Credits=30 \pe=1 \summry=(Bale|Action|251) From American Psycho to Batman Begins to Vice, Christian Bale is a bonafide A-list star But he missed out on plenty of huge roles along the way. >ab:xy_e0ra81-2 \Movie= Superman \actor= henry cavill \Genere=Action \Length=254 \Credits=28 \pe=1 \summry=(cavill|action|254) Henry William Dalgliesh Cavill is a British actor He is known for his portrayal of Charles Brandon in Showtime's The Tudors
- 从这些条目中提取
Name(>后的ID)和Length值,生成如下表格:
Name Length ab:xy_a0by98-2 234 ab:xy_c0ma65-1 251 ab:xy_e0ra81-2 254
之前用grep "pe=1" input.txt > output.txt只能提取到每个条目的第一行,无法获取后续描述,需要解决方法。
解决方案
1. 提取完整条目
用awk处理空行分隔的段落,匹配包含pe=1的段落并输出:
awk -v RS='' '/pe=1/' input.txt > full_entries.txt
-v RS='':将记录分隔符设为空,让awk把空行分隔的内容当作一个完整条目(段落)/pe=1/:匹配包含pe=1的条目,输出整个段落
执行后,full_entries.txt会包含所有符合要求的完整条目。
2. 生成指定表格
用awk提取所需字段,再用column -t对齐格式:
awk -v RS='' '/pe=1/ { match($0, /^>([^ ]+)/, name); match($0, /\\Length=([0-9]+)/, length); print name[1] "\t" length[1] }' input.txt | column -t > table.txt
match函数:分别提取>后的ID和\Length=后的数字column -t:自动对齐表格内容,让输出格式规整
执行后table.txt里的内容就是符合要求的表格:
Name Length ab:xy_a0by98-2 234 ab:xy_c0ma65-1 251 ab:xy_e0ra81-2 254
内容的提问来源于stack exchange,提问作者user2110417
相关产品推荐
相关产品推荐

