You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LXML XPath提取SEC财报数据丢失括号问题求助

Fixing Negative Value Extraction from SEC Filing Tables

Hey there, let's sort out this issue with extracting those parenthesized negative values from SEC filings. The problem here is that your original XPath isn't capturing the full text content (including the parentheses) nested inside the <a> tag within the target <td> elements.

What's Going Wrong?

Your current XPath //td[@class="nump"]/text() targets direct text nodes under the <td> element. But looking at the page structure you described:

<td class="num" ...> <a ..>(somevalue)</a></td>

The actual value (including parentheses for negatives) lives inside the <a> tag, not directly under the <td>. That's why you're only getting the numeric part without the parentheses—your XPath isn't reaching the right text node.

Solution Steps

  1. Adjust Your XPath to Grab Full Text
    Update your XPath to target the text inside the <a> tag within the num class <td>s. This will capture the complete content, including parentheses for negative values:

    //td[@class="nump"]/a/text()
    
  2. Process Raw Values to Convert Parentheses to Negatives
    Once you have the raw text (like (461,827) or 123,456), add a processing step to convert parenthesized values to negative numbers, and clean up commas for numeric conversion.

Modified Code Example

from lxml.html import fromstring
import requests

target_page = 'https://www.sec.gov/Archives/edgar/data/1564408/000156459017022434/R4.htm'
page = requests.get(target_page)
tree = fromstring(page.content)

# Grab full text (including parentheses) from the <a> tags inside target <td>s
raw_values = tree.xpath('//td[@class="nump"]/a/text()')

processed_values = []
for val in raw_values:
    cleaned_val = val.strip()  # Remove any leading/trailing whitespace
    if cleaned_val.startswith('(') and cleaned_val.endswith(')'):
        # Convert parenthesized value to negative number
        numeric_val = float(cleaned_val[1:-1].replace(',', '')) * -1
    else:
        # Convert positive value to float
        numeric_val = float(cleaned_val.replace(',', ''))
    processed_values.append(numeric_val)

# Check the results
print(processed_values)

This code will now correctly capture values like (461,827) as -461827.0 and positive values as their numeric equivalents.

内容的提问来源于stack exchange,提问作者Phil Dwan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:35:23