Python正则表达式无法从Outlook邮件提取关联文本与URL求助
问题排查:Outlook邮件关键词关联内容提取失败
需求说明
从Outlook指定文件夹的邮件中,先匹配邮件主题中的父关键词,再匹配正文内的子关键词,提取子关键词关联的段落文本与URL,但当前脚本仅能匹配关键词,无法提取对应关联内容。
示例邮件内容
01 事務用品・機器 大阪府警察大正警察署:指サック等の購入 :大阪市大正区 https://www.e-nyusatsu.pref.osaka.jp/CALS/Publish/EbController?Shori=SmallKokokuInfo&open_kokoku=01202350042214 01 事務用品・機器 府立学校大阪わかば高等学校:校内衛生用品7件 ★ :大阪市生野区 https://www.e-nyusatsu.pref.osaka.jp/CALS/Publish/EbController?Shori=SmallKokokuInfo&open_kokoku=01202350041978 01 事務用品・機器 府立学校工芸高等学校:イレパネ 他 購入 :大阪市阿倍野区 https://www.e-nyusatsu.pref.osaka.jp/CALS/Publish/EbController?Shori=SmallKokokuInfo&open_kokoku=01202350042117
配置文件(config2.json)
{ "folder_name": "調達プロジェクト", "output_file_path": "E:\\output", "output_file_name": "output.txt", "parent_keyword": "meeting", "child_keywords": ["土木一式工事", "産業用機器", "事務用品・機器"] }
现有Python代码
import win32com.client import os import json import logging import re def read_config(config_file): with open(config_file, 'r', encoding="utf-8") as f: config = json.load(f) return config def search_and_save_email(config): try: folder_name = config.get("folder_name", "") output_file_path = config.get("output_file_path", "") parent_keyword = config.get("parent_keyword", "") child_keywords = config.get("child_keywords", []) # Ensure the directory exists os.makedirs(output_file_path, exist_ok=True) outlook = win32com.client.Dispatch("Outlook.Application").GetNamespace("MAPI") inbox = outlook.GetDefaultFolder(6) # Find the user-created folder within the Inbox user_folder = None for folder in inbox.Folders: if folder.Name == folder_name: user_folder = folder break if user_folder is not None: # Search for emails with the parent keyword anywhere in the subject parent_keyword_pattern = re.compile(r'\b(?:' + '|'.join(map(re.escape, parent_keyword.split())) + r')\b', re.IGNORECASE) for item in user_folder.Items: if parent_keyword_pattern.findall(item.Subject): logging.info(f"Found parent keyword in Subject: {item.Subject}") # Parent keyword found, now search for child keywords in the body body_lower = item.Body.lower() # Initialize output_text outside the child keywords loop output_text = "" for child_keyword in child_keywords: # Search for child keyword in the body using regular expression child_keyword_pattern = re.compile(re.escape(child_keyword), re.IGNORECASE) matches = child_keyword_pattern.finditer(body_lower) for match in matches: logging.info(f"Found child keyword '{child_keyword}' at position {match.start()}-{match.end()}") # Extract the paragraph around the matched position paragraph_start = body_lower.rfind('\n', 0, match.start()) paragraph_end = body_lower.find('\n', match.end()) paragraph_text = item.Body[paragraph_start + 1:paragraph_end] # Extract URLs from the paragraph using a simple pattern url_pattern = re.compile(r'http[s]?://\S+') urls = url_pattern.findall(paragraph_text) # Append the results to the output_text output_text += f"Child Keyword: {child_keyword}\n" output_text += f"Paragraph Text: {paragraph_text}\n" output_text += f"URLs: {', '.join(urls)}\n\n" # Save the result to a text file output_file = os.path.join(output_file_path, f"{item.Subject.replace(' ', '_')}.txt") with open(output_file, 'w', encoding='utf-8') as f: f.write(output_text) logging.info(f"Saved results to {output_file}") else: logging.warning(f"Child keywords not found in folder '{folder_name}'.") else: logging.warning(f"Folder '{folder_name}' not found.") except Exception as e: logging.error(f"An error occurred: {str(e)}") if __name__ == "__main__": # Set up logging logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s') # Specify the path to the configuration file config_file_path = "E:\\config2.json" # Read configuration from the file config = read_config(config_file_path) # Search and save email based on the configuration search_and_save_email(config)
当前错误输出
Child Keyword: 土木一式工事 Paragraph Text: 土木一式工事 URLs: Child Keyword: 産業用機器 Paragraph Text: 19 産業用機器 URLs: Child Keyword: 産業用機器 Paragraph Text: 19 産業用機器 URLs:
问题原因
- 索引错位:将正文转成小写后匹配关键词,再用小写文本的索引去切割原始正文,可能导致位置偏移(虽日文无大小写,但特殊字符处理仍可能引发问题)。
- 段落范围过窄:当前仅提取关键词所在的单行,而示例中关联的机构信息和URL在关键词的后续行,未被纳入提取范围。
修复方案
修改后的代码
import win32com.client import os import json import logging import re def read_config(config_file): with open(config_file, 'r', encoding="utf-8") as f: config = json.load(f) return config def search_and_save_email(config): try: folder_name = config.get("folder_name", "") output_file_path = config.get("output_file_path", "") parent_keyword = config.get("parent_keyword", "") child_keywords = config.get("child_keywords", []) os.makedirs(output_file_path, exist_ok=True) outlook = win32com.client.Dispatch("Outlook.Application").GetNamespace("MAPI") inbox = outlook.GetDefaultFolder(6) user_folder = None for folder in inbox.Folders: if folder.Name == folder_name: user_folder = folder break if user_folder is not None: # 父关键词匹配逻辑不变 parent_keyword_pattern = re.compile(r'\b(?:' + '|'.join(map(re.escape, parent_keyword.split())) + r')\b', re.IGNORECASE) for item in user_folder.Items: if parent_keyword_pattern.findall(item.Subject): logging.info(f"Found parent keyword in Subject: {item.Subject}") body = item.Body output_text = "" # 预编译所有子关键词的正则,用于查找下一个关键词位置 all_child_pattern = re.compile('|'.join(map(re.escape, child_keywords)), re.IGNORECASE) for child_keyword in child_keywords: child_pattern = re.compile(re.escape(child_keyword), re.IGNORECASE) matches = child_pattern.finditer(body) for match in matches: logging.info(f"Found child keyword '{child_keyword}' at position {match.start()}-{match.end()}") # 提取段落起始:关键词所在行的开头 paragraph_start = body.rfind('\n', 0, match.start()) + 1 # 查找下一个子关键词的位置,作为段落结束;如果没有则取正文末尾 next_match = all_child_pattern.search(body, match.end()) paragraph_end = next_match.start() if next_match else len(body) # 提取整个关联块,包含后续内容和URL paragraph_text = body[paragraph_start:paragraph_end].strip() # 提取该块内的所有URL url_pattern = re.compile(r'http[s]?://\S+') urls = url_pattern.findall(paragraph_text) output_text += f"Child Keyword: {child_keyword}\n" output_text += f"Paragraph Text: {paragraph_text}\n" output_text += f"URLs: {', '.join(urls)}\n\n" output_file = os.path.join(output_file_path, f"{item.Subject.replace(' ', '_')}.txt") with open(output_file, 'w', encoding='utf-8') as f: f.write(output_text) logging.info(f"Saved results to {output_file}") else: logging.warning(f"No parent keyword found in email subject: {item.Subject}") else: logging.warning(f"Folder '{folder_name}' not found.") except Exception as e: logging.error(f"An error occurred: {str(e)}") if __name__ == "__main__": logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s') config_file_path = "E:\\config2.json" config = read_config(config_file_path) search_and_save_email(config)
关键修改点
- 移除小写转换:直接在原始正文上进行匹配,避免索引错位问题。
- 扩展段落范围:找到子关键词后,向后查找下一个子关键词的位置作为段落结束,确保包含关联的机构信息和URL。
- 优化日志提示:修正原日志中错误的提示信息(原提示"Child keywords not found"实际是父关键词未匹配)。
预期输出示例
Child Keyword: 事務用品・機器 Paragraph Text: 01 事務用品・機器 大阪府警察大正警察署:指サック等の購入 :大阪市大正区 https://www.e-nyusatsu.pref.osaka.jp/CALS/Publish/EbController?Shori=SmallKokokuInfo&open_kokoku=01202350042214 URLs: https://www.e-nyusatsu.pref.osaka.jp/CALS/Publish/EbController?Shori=SmallKokokuInfo&open_kokoku=01202350042214 Child Keyword: 事務用品・機器 Paragraph Text: 01 事務用品・機器 府立学校大阪わかば高等学校:校内衛生用品7件 ★ :大阪市生野区 https://www.e-nyusatsu.pref.osaka.jp/CALS/Publish/EbController?Shori=SmallKokokuInfo&open_kokoku=01202350041978 URLs: https://www.e-nyusatsu.pref.osaka.jp/CALS/Publish/EbController?Shori=SmallKokokuInfo&open_kokoku=01202350041978
内容的提问来源于stack exchange,提问作者Azmol Hossain
相关产品推荐
相关产品推荐

