38 lines
5.4 KiB
Markdown
38 lines
5.4 KiB
Markdown
# Keyed factor diagnostics, candidate API 0.1.0
|
|
|
|
`quant_engine.factor_diagnostics` is a pure calculation module for caller-supplied observations. It does not change the older `factor_library.ic_summary` API or the governed factor-set contracts. Its calculation/API version does not establish a production algorithm, source qualification, historical availability or execution eligibility.
|
|
|
|
## Inputs and outputs
|
|
|
|
Factor values are a DataFrame indexed by unique `(date, asset)` keys, with unique factor columns. Dates must be a naive DatetimeIndex of midnight session labels; strings and timezone-aware/intraday timestamps are rejected. Order may vary and is normalized without modifying inputs. Identifiers must be nonempty printable strings. Values are real finite numbers or missing (`None`, `pd.NA`, NaN); booleans, numeric strings, infinities and duplicate keys are rejected.
|
|
|
|
`daily_ic(factors, returns, min_pairs=3, min_days=2)` aligns the forward-return Series by complete keys inside each date. The factor keys define the observed universe. Extra return keys are ignored and counted, while absent return keys remain missing. Every observed factor/date survives, including zero valid pairs. Output includes daily Pearson and average-tie Spearman values, actual pair/observation/value counts and separate reasons. Returns must already be labels for the intended forward interval; this function does not infer their unit, calendar, origin or availability.
|
|
|
|
Daily summaries weight each valid date equally. Mean and sample standard deviation (`ddof=1`) use valid daily ICs, not the number of securities. IR is unannualized `mean/std`; the descriptive IID statistic is `IR*sqrt(valid_days)`, with two-sided Student-t p and `df=valid_days-1`. Pearson and RankIC have separate valid/missing day counts. No-valid-day and insufficient-day results remain explicit. Constant series have standard deviation zero and undefined ratios.
|
|
|
|
The absolute standard-deviation resolution for ratio statistics is `32*float64 epsilon` (about 7.11e-15) because IC is bounded to [-1,1]. Nonconstant dispersion at or below that resolution retains its mean, observed standard deviation and counts, but IR/t/p are null with `below_resolution`. This candidate numerical policy prevents floating-point noise between equivalent cross sections from becoming extreme significance. It is not a statistical materiality threshold. Serial correlation, overlapping forward intervals and effective sample size are **not corrected**; t/p do not establish inferential validity or decision admission.
|
|
|
|
`correlation_matrix(factors, method='pearson', min_pairs=3)` pools complete `(date,asset)` pairs for each cell. It is not the mean of daily cross-sectional correlations: dates with more pairs contribute more observations. Every cell reports pair count and status. Empty, short or constant diagonals are null, not identity values. Pairwise deletion can produce a non-positive-semidefinite matrix; this is not an admitted risk/covariance matrix. Pearson translates before scaling to preserve representable small differences near a large offset, with scale-first fallback only if subtraction overflows. Spearman ranks original paired observations to avoid creating ties through underflow.
|
|
|
|
## Explicit forward-return intervals
|
|
|
|
`forward_returns(prices, sessions=..., entry_lag_sessions=..., holding_sessions=..., price_field=..., price_basis=...)` accepts a keyed price Series and an explicit, unique, increasing session calendar. Lag must be an integer >=0 and holding an integer >=1; booleans are rejected. The returned `ForwardReturns` object owns a return Series and interval DataFrame for every supplied session and observed asset, plus method metadata.
|
|
|
|
For signal session `t`, entry is session `t+lag`, exit is `t+lag+holding`, and the label is `P_exit/P_entry-1`. Endpoint prices must be positive when present and comparable under the caller-declared field/basis. Missing endpoint prices yield `missing_price`; calendar-tail insufficiency yields `insufficient_calendar`; nonfinite arithmetic yields `numerical_failure`. No filling, per-asset dropna calendar, next-available-price jump or daily-return summation is used. Intermediate prices are not required for this endpoint ratio. A lag of zero describes a same-session price basis and does not mean a closing signal can trade at the same close.
|
|
|
|
## Reuse decision and verification
|
|
|
|
需求:按完整证券/日期键提供逐日IC、相关矩阵与显式前瞻标签,保留样本及未定义原因。
|
|
|
|
已有方案:`factor_library.ic_summary(periods=(1,))`的单截面Pearson和`spearman_ic`;旧多周期rolling-sum不符合显式端点区间,旧摘要也不代替逐日统计。
|
|
|
|
候选开源方案:不需要,已有pandas/numpy/scipy及核心函数足够,无新增依赖。
|
|
|
|
推荐方案:在核心中二次封装现有单截面相关函数,增加观察键、分组、计数与区间合同。
|
|
|
|
原因:保持计算归属quant_engine,平台只做输入/输出适配;旧API保持兼容。
|
|
|
|
风险:浮点分辨率、缺失机制、序列依赖和PIT均须显式记录,纯合成通过不能升级正式资格。
|
|
|
|
Tests include manually checkable daily `[1,-0.5,0]` ICs with three effective dates, pairwise missingness, empty/constant results, average ranks, extreme numeric ranges, affine-equivalent daily ICs, large-offset Pearson precision, and explicit-calendar endpoint labels. Existing factor-library tests remain unchanged. No database, provider, production recomputation or ETL is part of this contract.
|