You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多XML文件无解析合并下载脚本执行后输出为空问题排查

问题排查:合并XML URL内容但输出文件为空

问题描述

我需要将多个独立URL的XML文件不经过解析直接合并到单个文件中用于关键词分析,编写了如下Python脚本,脚本能正常执行,但输出文件里没有任何内容,请问问题出在哪里?

import requests

# List of URLs of the XML pages you want to scrape
xml_urls = [
    "https://nsearchives.nseindia.com/corporate/xbrl/CG_92090_946801_11102023020327_WEB.xml",
    "https://nsearchives.nseindia.com/corporate/xbrl/CG_92138_947508_11102023050314_WEB.xml",
    # Add more URLs as needed
]

# Output file where the content will be saved
output_file = "output.txt"

# Function to scrape XML content from a URL and save it to the output file
def scrape_and_save_xml(url, output_file):
    try:
        response = requests.get(url)
        if response.status_code == 200:
            # Write the raw content to the output file
            with open(output_file, "a", encoding="utf-8") as file:
                file.write(response.text + "\n")
            print(f"Scraped {url}")
        else:
            print(f"Failed to fetch {url} (Status Code: {response.status_code})")
    except Exception as e:
        print(f"Error while scraping {url}: {str(e)}")

# Clear the contents of the output file (if it already exists)
with open(output_file, "w", encoding="utf-8") as file:
    file.truncate()

# Loop through the XML URLs and scrape their content
for xml_url in xml_urls:
    scrape_and_save_xml(xml_url, output_file)

print("Scraping completed. The content has been saved to", output_file)

问题原因与解决方案

1. 反爬机制拦截请求

NSE India的服务器会验证请求的User-Agent头,默认requests.get()不带该头,会被服务器拒绝(返回空内容或403状态码)。需要给请求添加模拟浏览器的请求头。

2. 未验证响应实际内容

即使响应状态码是200,也可能返回反爬页面而非目标XML内容,建议添加内容校验步骤。

3. 文件操作冗余

清空文件的代码可以简化,使用"w"模式打开文件时会自动清空原有内容,无需额外调用file.truncate()。

修改后的脚本

import requests

# List of URLs of the XML pages you want to scrape
xml_urls = [
    "https://nsearchives.nseindia.com/corporate/xbrl/CG_92090_946801_11102023020327_WEB.xml",
    "https://nsearchives.nseindia.com/corporate/xbrl/CG_92138_947508_11102023050314_WEB.xml",
    # Add more URLs as needed
]

# Output file where the content will be saved
output_file = "output.txt"

# 添加模拟浏览器的请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

# Function to scrape XML content from a URL and save it to the output file
def scrape_and_save_xml(url, output_file):
    try:
        response = requests.get(url, headers=headers)
        response.raise_for_status()  # 主动抛出HTTP错误
        
        # 验证是否获取到XML内容
        if "<?xml" not in response.text[:100]:
            print(f"Warning: {url} 返回的内容不是XML,可能被反爬拦截")
            print("响应内容预览:", response.text[:500])
            return
        
        # Write the raw content to the output file
        with open(output_file, "a", encoding="utf-8") as file:
            file.write(response.text + "\n")
        print(f"Scraped {url}")
    except requests.exceptions.HTTPError as e:
        print(f"Failed to fetch {url} (HTTP Error: {str(e)})")
    except Exception as e:
        print(f"Error while scraping {url}: {str(e)}")

# 清空文件("w"模式自动清空)
with open(output_file, "w", encoding="utf-8") as file:
    pass

# Loop through the XML URLs and scrape their content
for xml_url in xml_urls:
    scrape_and_save_xml(xml_url, output_file)

print("Scraping completed. The content has been saved to", output_file)

内容的提问来源于stack exchange,提问作者Chetan Lunkar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 19:05:21