You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取SEC EDGAR URL中的特定数字段?

Methods to Extract Target Numbers from SEC EDGAR URLs

The numbers you're targeting are the unique identifiers located right after /edgar/data/ in each URL. Here are a few straightforward ways to extract them:

1. Using Python with Regular Expressions

This is a flexible approach if you’re comfortable with basic coding. You can use Python’s built-in re module to match the pattern and pull out the numbers.

import re

urls = [
    "https://www.sec.gov/Archives/edgar/data/1002638/000100263816000080/exhibit211subsidiarylisting.htm",
    "http://www.sec.gov/Archives/edgar/data/1013871/000101387113000003/exhibit21110k2012.htm",
    "http://www.sec.gov/Archives/edgar/data/1420800/000142080014000006/exhibit211subsidiariesofth.htm",
    "http://www.sec.gov/Archives/edgar/data/1305014/000130501415000119/a9302015exhibit21.htm"
]

# Regex pattern to capture the number after /edgar/data/
pattern = r"/edgar/data/(\d+)/"
extracted_numbers = [re.search(pattern, url).group(1) for url in urls]

# Output the numbers in the desired format
print(" ".join(extracted_numbers))

Running this code will output exactly what you need: 1002638 1013871 1420800 1305014

2. Using Command Line Tools (Grep)

If you prefer working in a terminal, save your URLs to a text file (e.g., urls.txt) and use this grep command to extract the numbers:

grep -oP '/edgar/data/\K\d+' urls.txt | tr '\n' ' '
  • -o tells grep to output only the matched parts
  • -P enables Perl-compatible regex
  • \K discards the preceding text (/edgar/data/) so only the number is returned
  • tr '\n' ' ' converts newlines to spaces for the space-separated format

3. Using Text Editors with Regex Support

Most modern text editors (VS Code, Sublime Text, Notepad++) let you use regex to extract these numbers quickly:

  • Open the file with your URLs
  • Open the find dialog (usually Ctrl+F or Cmd+F)
  • Enable regex mode (look for a .* icon or check "Use Regular Expression")
  • Enter the pattern /edgar/data/(\d+)/
  • Use the editor’s "Find All" feature to collect all matches, then copy the captured numeric groups

内容的提问来源于stack exchange,提问作者Gautam Biswas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:51:07