从现有NumPy数组派生子数组计算相关系数,找出相关性最强的两只指数
Got it, let's walk through solving this problem step by step. You've got a NumPy array with ticker, currency, date, and exchange rate data, and you want to figure out which two indices have the strongest correlation in their exchange rate values. We'll use a mix of NumPy and pandas here—pandas makes reshaping the data into a correlation-friendly format much simpler.
Step 1: Import Required Libraries
First, make sure you have these installed (if not, run pip install numpy pandas), then import them:
import numpy as np import pandas as pd
Step 2: Convert the NumPy Array to a Pandas DataFrame
Turning your raw array into a DataFrame will let us easily manipulate the structure:
# Your original array (truncated example) data = np.array([ ['AIndex', 'AUD', '2018-01-01', 1.28], ['BIndex', 'EUR', '2018-01-01', -0.248], ['CIndex', 'HKD', '2018-01-01', -1.18], ['DIndex', 'JPY', '2018-04-06', -1.046], ['EIndex', 'EUR', '2018-04-06', 1.092], ['FIndex', 'CAD', '2018-04-06', 0.61] ], dtype=object) # Create DataFrame with clear column names df = pd.DataFrame(data, columns=['Ticker', 'Currency', 'Date', 'ExchangeRate']) # Convert exchange rate values to numeric type (they're stored as objects initially) df['ExchangeRate'] = pd.to_numeric(df['ExchangeRate'])
Step 3: Reshape Data into a Correlation-Friendly Format
We need a wide-format DataFrame where each column is an index, each row is a date, and the values are the exchange rates. This alignment is critical for meaningful correlation calculations:
# Pivot the data: rows = dates, columns = tickers, values = exchange rates pivoted_df = df.pivot(index='Date', columns='Ticker', values='ExchangeRate') # Optional: If some indices have missing dates, fill gaps (example using forward fill) # pivoted_df = pivoted_df.fillna(method='ffill')
Step 4: Calculate the Correlation Matrix
Now compute the Pearson correlation coefficient matrix—this gives us the correlation between every pair of indices:
corr_matrix = pivoted_df.corr()
Step 5: Find the Pair with the Strongest Correlation
We need to ignore the diagonal (since an index is perfectly correlated with itself) and find the maximum absolute correlation value (to capture both strong positive and negative correlations):
# Mask the diagonal (values of 1) to exclude self-correlation mask = np.triu(np.ones_like(corr_matrix, dtype=bool)) corr_matrix_masked = corr_matrix.mask(mask) # Find the maximum absolute correlation value max_corr_value = corr_matrix_masked.abs().max().max() # Get the corresponding pair of tickers top_corr_pair = corr_matrix_masked.abs().stack().idxmax() print(f"The pair of indices with the strongest correlation is {top_corr_pair}, with a correlation coefficient of {max_corr_value:.4f}")
Quick Notes
- The pivot step ensures all exchange rates are aligned by date—without this, correlation calculations would be meaningless.
- Using absolute value lets us catch strong inverse correlations (e.g., if one index's rate rises when another falls) just as easily as positive ones.
- If your dataset has missing values, choose a fill method that makes sense for your data (forward fill, mean, or interpolation) before calculating correlations.
内容的提问来源于stack exchange,提问作者dirtyw0lf

