多XML文件无解析合并下载脚本执行后输出为空问题排查
问题排查:合并XML URL内容但输出文件为空
问题描述
我需要将多个独立URL的XML文件不经过解析直接合并到单个文件中用于关键词分析,编写了如下Python脚本,脚本能正常执行,但输出文件里没有任何内容,请问问题出在哪里?
import requests # List of URLs of the XML pages you want to scrape xml_urls = [ "https://nsearchives.nseindia.com/corporate/xbrl/CG_92090_946801_11102023020327_WEB.xml", "https://nsearchives.nseindia.com/corporate/xbrl/CG_92138_947508_11102023050314_WEB.xml", # Add more URLs as needed ] # Output file where the content will be saved output_file = "output.txt" # Function to scrape XML content from a URL and save it to the output file def scrape_and_save_xml(url, output_file): try: response = requests.get(url) if response.status_code == 200: # Write the raw content to the output file with open(output_file, "a", encoding="utf-8") as file: file.write(response.text + "\n") print(f"Scraped {url}") else: print(f"Failed to fetch {url} (Status Code: {response.status_code})") except Exception as e: print(f"Error while scraping {url}: {str(e)}") # Clear the contents of the output file (if it already exists) with open(output_file, "w", encoding="utf-8") as file: file.truncate() # Loop through the XML URLs and scrape their content for xml_url in xml_urls: scrape_and_save_xml(xml_url, output_file) print("Scraping completed. The content has been saved to", output_file)
问题原因与解决方案
1. 反爬机制拦截请求
NSE India的服务器会验证请求的User-Agent头,默认requests.get()不带该头,会被服务器拒绝(返回空内容或403状态码)。需要给请求添加模拟浏览器的请求头。
2. 未验证响应实际内容
即使响应状态码是200,也可能返回反爬页面而非目标XML内容,建议添加内容校验步骤。
3. 文件操作冗余
清空文件的代码可以简化,使用"w"模式打开文件时会自动清空原有内容,无需额外调用file.truncate()。
修改后的脚本
import requests # List of URLs of the XML pages you want to scrape xml_urls = [ "https://nsearchives.nseindia.com/corporate/xbrl/CG_92090_946801_11102023020327_WEB.xml", "https://nsearchives.nseindia.com/corporate/xbrl/CG_92138_947508_11102023050314_WEB.xml", # Add more URLs as needed ] # Output file where the content will be saved output_file = "output.txt" # 添加模拟浏览器的请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } # Function to scrape XML content from a URL and save it to the output file def scrape_and_save_xml(url, output_file): try: response = requests.get(url, headers=headers) response.raise_for_status() # 主动抛出HTTP错误 # 验证是否获取到XML内容 if "<?xml" not in response.text[:100]: print(f"Warning: {url} 返回的内容不是XML,可能被反爬拦截") print("响应内容预览:", response.text[:500]) return # Write the raw content to the output file with open(output_file, "a", encoding="utf-8") as file: file.write(response.text + "\n") print(f"Scraped {url}") except requests.exceptions.HTTPError as e: print(f"Failed to fetch {url} (HTTP Error: {str(e)})") except Exception as e: print(f"Error while scraping {url}: {str(e)}") # 清空文件("w"模式自动清空) with open(output_file, "w", encoding="utf-8") as file: pass # Loop through the XML URLs and scrape their content for xml_url in xml_urls: scrape_and_save_xml(xml_url, output_file) print("Scraping completed. The content has been saved to", output_file)
内容的提问来源于stack exchange,提问作者Chetan Lunkar
相关产品推荐
相关产品推荐

