如何将Factiva文本文件批量解析为Pandas DataFrame?
问题:Factiva文本转Pandas DataFrame解析方案
本人是Python编程新手,希望将Factiva文本文件转换为Pandas DataFrame。以下为Factiva文本样例:
Section US Treas Dept on Feb 13 makes public text of speech made by Sec Shultz at... 228 words 15 February 1974 New York Times Abstracts NYTA Pg. 1, Col. 1 English c. 1974 New York Times Company US Treas Dept on Feb 13 makes public text of speech made by Sec Shultz at closed session of 13-nation energy conf in which he lists ideas and suggestions for dealing with purely financial problems raised for world by huge increase in price of oil. High officials elaborate at background news briefing. Shultz emphasizes that there is no internatl financial arrangement which can offset real effects of oil pricechanges. Adds that there is no way to print money and use it to 'paper over' problem of high price of oil and its impact on employment, prices and real incomes in rich and poor countries. Shultz's ideas were designed to keep real problem from being made worse by huge new money flows involved, amounting to some $50-billion more in '74 than in '73 from oil-consuming countries to oil-producing countries. Suggestions include possible coordination of borrowings by indus nations on world financial mkts that will be needed to cover balance-of-payments deficits, new ways of aiding 'poorest' of poor countries, which will be unable to borrow, possible new 'investment fund' to help oil countries place long-term investments, new reciprocal currency 'swap' arrangements among leading central banks and others. Shultz suggestions for industrial countries outlined (M). Document nyta000020011127d62f07et9我期望提取的字段包括section、word count、date、news source、language以及content,且需支持循环处理数百个Factiva文本文件。以下为我目前编写的Python代码:
import pandas as pd # Sample text content text_content = """ Section US Treas Dept on Feb 13 makes public text of speech made by Sec Shultz at... 228 words 15 February 1974 New York Times Abstracts NYTA Pg. 1, Col. 1 English c. 1974 New York Times Company US Treas Dept on Feb 13 makes public text of speech made by Sec Shultz at closed session of 13-nation energy conf in which he lists ideas and suggestions for dealing with purely financial problems raised for world by huge increase in price of oil. High officials elaborate at background news briefing. Shultz emphasizes that there is no internatl financial arrangement which can offset real effects of oil price changes. Adds that there is no way to print money and use it to 'paper over' problem of high price of oil and its impact on employment, prices and real incomes in rich and poor countries. Shultz's ideas were designed to keep real problem from being made worse by huge new money flows involved, amounting to some $50-billion more in '74 than in '73 from oil-consuming countries to oil-producing countries. Suggestions include possible coordination of borrowings by indus nations on world financial mkts that will be needed to cover balance-of-payments deficits, new ways of aiding 'poorest' of poor countries, which will be unable to borrow, possible new 'investment fund' to help oil countries place long-term investments, new reciprocal currency 'swap' arrangements among leading central banks and others. Shultz suggestions for industrial countries outlined (M). Document nyta000020011127d62f07et9 """ # Define a list of headers headers = ["Section", "words", "Date", "Source"] # Create a list to store dictionaries for each section data_list = [] # Split the text based on the headers for i in range(len(headers)-1): header_start = headers[i] header_end = headers[i+1] section_content = text_content.split(header_start)[1].split(header_end)[0].strip() # Create a dictionary for each section section_dict = { 'Header': header_start, 'Content': section_content } data_list.append(section_dict) # Create a pandas DataFrame from the list of dictionaries df = pd.DataFrame(data_list) # Display the DataFrame print(df)请指导我如何实现正确的解析逻辑。
解决方案
解析思路
根据Factiva文本的固定格式,采用以下逻辑提取目标字段:
- 按空行分割文本,得到独立的内容块
- 从对应位置的块中提取section、元数据(词数、日期、来源、语言)和正文
- 将单文件解析结果封装为字典,批量处理时存入列表后转换为DataFrame
完整代码实现
import pandas as pd import os def parse_factiva_text(text): # 按空行分割文本,过滤无效空块 blocks = [block.strip() for block in text.split('\n\n') if block.strip()] # 提取section:移除第一个块的"Section"前缀 section = blocks[0].replace('Section', '').strip() # 提取元数据:第二个块按换行拆分后取对应位置内容 meta_lines = blocks[1].split('\n') word_count = int(meta_lines[0].replace(' words', '')) date = meta_lines[1] news_source = meta_lines[2] language = meta_lines[5] # 提取正文:从第三个块到倒数第二个块(跳过最后一个Document标识块) content = '\n'.join(blocks[2:-1]) return { 'section': section, 'word_count': word_count, 'date': date, 'news_source': news_source, 'language': language, 'content': content } def batch_process_factiva_files(folder_path): all_data = [] # 遍历目标文件夹下的所有文件 for filename in os.listdir(folder_path): # 可根据实际文件格式调整后缀判断 if filename.endswith('.txt'): file_path = os.path.join(folder_path, filename) with open(file_path, 'r', encoding='utf-8') as f: text = f.read() try: parsed_data = parse_factiva_text(text) # 可选:添加文件名字段便于溯源 parsed_data['filename'] = filename all_data.append(parsed_data) except Exception as e: print(f"处理文件 {filename} 出错:{str(e)}") # 转换为DataFrame并返回 return pd.DataFrame(all_data) # 测试与使用示例 if __name__ == '__main__': # 单文件解析测试 sample_text = """ Section US Treas Dept on Feb 13 makes public text of speech made by Sec Shultz at... 228 words 15 February 1974 New York Times Abstracts NYTA Pg. 1, Col. 1 English c. 1974 New York Times Company US Treas Dept on Feb 13 makes public text of speech made by Sec Shultz at closed session of 13-nation energy conf in which he lists ideas and suggestions for dealing with purely financial problems raised for world by huge increase in price of oil. High officials elaborate at background news briefing. Shultz emphasizes that there is no internatl financial arrangement which can offset real effects of oil price changes. Adds that there is no way to print money and use it to 'paper over' problem of high price of oil and its impact on employment, prices and real incomes in rich and poor countries. Shultz's ideas were designed to keep real problem from being made worse by huge new money flows involved, amounting to some $50-billion more in '74 than in '73 from oil-consuming countries to oil-producing countries. Suggestions include possible coordination of borrowings by indus nations on world financial mkts that will be needed to cover balance-of-payments deficits, new ways of aiding 'poorest' of poor countries, which will be unable to borrow, possible new 'investment fund' to help oil countries place long-term investments, new reciprocal currency 'swap' arrangements among leading central banks and others. Shultz suggestions for industrial countries outlined (M). Document nyta000020011127d62f07et9 """ single_result = parse_factiva_text(sample_text) print("单文件解析结果:") print(pd.DataFrame([single_result])) # 批量处理示例(替换为你的Factiva文件所在文件夹路径) # df = batch_process_factiva_files('./factiva_docs') # df.to_csv('factiva_parsed_result.csv', index=False) # print("批量处理完成,结果已保存为factiva_parsed_result.csv")
代码说明
parse_factiva_text:负责单文件的核心解析逻辑,针对Factiva固定格式提取指定字段batch_process_factiva_files:遍历目标文件夹,批量处理所有符合格式的文件,加入异常捕获避免单个文件出错中断流程- 可根据实际文件类型调整
filename.endswith('.txt')的判断条件,确保覆盖所有Factiva文件
内容的提问来源于stack exchange,提问作者Aris Zoleta
相关产品推荐
相关产品推荐

