feat: add keyed daily factor diagnostics and forward intervals
CI / lite (pull_request) Canceled after 0s

This commit is contained in:
ao gong
2026-10-04 03:47:02 +08:00
parent 861c1e97a8
commit 69383a30b0
5 changed files with 692 additions and 0 deletions
+1
View File
@@ -32,6 +32,7 @@
- `retrospective_*_contracts` — 未发布的显式 v2 回顾性合同:区分历史业务日期与实际可得/计算时间,保留 v1 和现有金融公式,不授予历史可得性、发布或执行权限;见 [v2 接口说明](docs/RETROSPECTIVE_COMPUTATION_V2.md)
- `attribution` — 基于实际成交后持仓的隔夜 / 日内 / 交易成本逐日收益归因与闭合审计
- `metrics` — 绝对绩效 + 严格日期对齐的 TE / IR / alpha / beta 基准相对绩效
- `factor_diagnostics` — 候选0.1.0:完整键配对、逐日IC/RankIC、样本与未定义值、显式日历前瞻标签;见 [诊断合同](docs/factor-diagnostics.md)
- `factor_library` — 通用方法(turnover / winsorize / IC / OLS / jb_test)
- `portfolio_decomp` — 组合分解(risk_parity / mean_variance / 因子归因)
- `risk` — ndarray 低层风险公式 + 标签安全、可分组的 Euler 成分风险分解
+37
View File
@@ -0,0 +1,37 @@
# Keyed factor diagnostics, candidate API 0.1.0
`quant_engine.factor_diagnostics` is a pure calculation module for caller-supplied observations. It does not change the older `factor_library.ic_summary` API or the governed factor-set contracts. Its calculation/API version does not establish a production algorithm, source qualification, historical availability or execution eligibility.
## Inputs and outputs
Factor values are a DataFrame indexed by unique `(date, asset)` keys, with unique factor columns. Dates must be a naive DatetimeIndex of midnight session labels; strings and timezone-aware/intraday timestamps are rejected. Order may vary and is normalized without modifying inputs. Identifiers must be nonempty printable strings. Values are real finite numbers or missing (`None`, `pd.NA`, NaN); booleans, numeric strings, infinities and duplicate keys are rejected.
`daily_ic(factors, returns, min_pairs=3, min_days=2)` aligns the forward-return Series by complete keys inside each date. The factor keys define the observed universe. Extra return keys are ignored and counted, while absent return keys remain missing. Every observed factor/date survives, including zero valid pairs. Output includes daily Pearson and average-tie Spearman values, actual pair/observation/value counts and separate reasons. Returns must already be labels for the intended forward interval; this function does not infer their unit, calendar, origin or availability.
Daily summaries weight each valid date equally. Mean and sample standard deviation (`ddof=1`) use valid daily ICs, not the number of securities. IR is unannualized `mean/std`; the descriptive IID statistic is `IR*sqrt(valid_days)`, with two-sided Student-t p and `df=valid_days-1`. Pearson and RankIC have separate valid/missing day counts. No-valid-day and insufficient-day results remain explicit. Constant series have standard deviation zero and undefined ratios.
The absolute standard-deviation resolution for ratio statistics is `32*float64 epsilon` (about 7.11e-15) because IC is bounded to [-1,1]. Nonconstant dispersion at or below that resolution retains its mean, observed standard deviation and counts, but IR/t/p are null with `below_resolution`. This candidate numerical policy prevents floating-point noise between equivalent cross sections from becoming extreme significance. It is not a statistical materiality threshold. Serial correlation, overlapping forward intervals and effective sample size are **not corrected**; t/p do not establish inferential validity or decision admission.
`correlation_matrix(factors, method='pearson', min_pairs=3)` pools complete `(date,asset)` pairs for each cell. It is not the mean of daily cross-sectional correlations: dates with more pairs contribute more observations. Every cell reports pair count and status. Empty, short or constant diagonals are null, not identity values. Pairwise deletion can produce a non-positive-semidefinite matrix; this is not an admitted risk/covariance matrix. Pearson translates before scaling to preserve representable small differences near a large offset, with scale-first fallback only if subtraction overflows. Spearman ranks original paired observations to avoid creating ties through underflow.
## Explicit forward-return intervals
`forward_returns(prices, sessions=..., entry_lag_sessions=..., holding_sessions=..., price_field=..., price_basis=...)` accepts a keyed price Series and an explicit, unique, increasing session calendar. Lag must be an integer >=0 and holding an integer >=1; booleans are rejected. The returned `ForwardReturns` object owns a return Series and interval DataFrame for every supplied session and observed asset, plus method metadata.
For signal session `t`, entry is session `t+lag`, exit is `t+lag+holding`, and the label is `P_exit/P_entry-1`. Endpoint prices must be positive when present and comparable under the caller-declared field/basis. Missing endpoint prices yield `missing_price`; calendar-tail insufficiency yields `insufficient_calendar`; nonfinite arithmetic yields `numerical_failure`. No filling, per-asset dropna calendar, next-available-price jump or daily-return summation is used. Intermediate prices are not required for this endpoint ratio. A lag of zero describes a same-session price basis and does not mean a closing signal can trade at the same close.
## Reuse decision and verification
需求:按完整证券/日期键提供逐日IC、相关矩阵与显式前瞻标签,保留样本及未定义原因。
已有方案:`factor_library.ic_summary(periods=(1,))`的单截面Pearson和`spearman_ic`;旧多周期rolling-sum不符合显式端点区间,旧摘要也不代替逐日统计。
候选开源方案:不需要,已有pandas/numpy/scipy及核心函数足够,无新增依赖。
推荐方案:在核心中二次封装现有单截面相关函数,增加观察键、分组、计数与区间合同。
原因:保持计算归属quant_engine,平台只做输入/输出适配;旧API保持兼容。
风险:浮点分辨率、缺失机制、序列依赖和PIT均须显式记录,纯合成通过不能升级正式资格。
Tests include manually checkable daily `[1,-0.5,0]` ICs with three effective dates, pairwise missingness, empty/constant results, average ranks, extreme numeric ranges, affine-equivalent daily ICs, large-offset Pearson precision, and explicit-calendar endpoint labels. Existing factor-library tests remain unchanged. No database, provider, production recomputation or ETL is part of this contract.
+11
View File
@@ -0,0 +1,11 @@
# Quant OS factor diagnostics core increment
Delivery: quant-os-factor-diagnostics-20261004; branch codex/quant-os-factor-diagnostics-20261004. Base is remote-confirmed origin/main 861c1e97a8bf1c3e962c4cd1ee88ef58e6b9ddd5. This is the independent core-repository increment consumed by the existing platform delivery/Draft #102, not a new platform branch or chat. Root in chat 01a0bc8f-dcaf-7452-9ab3-3215df6dfa97 is the sole core code/Git writer; reviewers are read-only. Primary, merged retrospective-v2 and old Alpha158 worktrees and their user files remain untouched.
The fixed framework source is 16e96351fbc5bd5a918c7deff556490bbc44fb99. Valid source cleanliness/version evidence was reused; actual core feature/worktree router ranges and applicable AGENTS were read. Resume/review model-policy critical gpt-6-astra/xhigh passed; actual runtime remains unknown. The core has no stage ledger at its declared path, so no foreign stage state is borrowed. Lifecycle created this task after policy-check; old completed trees were retained for ignored local data, no old lease/state was rewritten. The new .venv was created through lifecycle run with the existing frozen lock and dev extras, offline from local caches.
Scope is L3 mathematical behavior, pure memory/caller data. The new candidate module provides daily IC/RankIC summaries, keyed pooled correlation matrices and explicit-calendar forward-return labels. Production algorithm/source/PIT/execution qualifications remain unestablished and decision_eligible=false. Existing factor-library and governed factor-set contracts are unchanged. Reuse rationale and exact numerical/statistical assumptions are in docs/factor-diagnostics.md.
Evidence so far: initial47 tests failed for the missing API; first implementation46 passed and one strict floating-zero assertion failed, corrected to a stated1e-15 tolerance without clipping values. A genuine extreme-rank scaling regression was reproduced and fixed by ranking original values. An affine-equivalent IC case reproduced enormous IR/t from rounding dispersion, fixed with the public32eps resolution policy. Independent reviewer found two representable-offset Pearson errors; both were reproduced and fixed by translate-before-scale. Latest54 new plus16 old tests=70 passed; final whole-core run1093 passed in12.82s with1166 existing pandas deprecation warnings. Targeted Ruff and final module-only mypy passed. Reviewer rechecked the two numeric fixes with tiny independent samples, no new blocker.
This remains an unshipped candidate: document and full applicable checks, stable revision binding, platform thin consumption and synthetic UI evidence must be completed before final delivery gates. No source/provider/NAS query, production ETL, migration, release, deployment or trade has been authorized by this increment. Reuse this same lifecycle/branch/PR across further windows; do not treat a push as final delivery.
+325
View File
@@ -0,0 +1,325 @@
"""Pure, observation-keyed factor diagnostics on caller-supplied data.
These low-level candidate calculations neither fetch sources nor grant data,
historical-availability, production-algorithm or execution qualification. The
legacy factor_library API remains unchanged. IC summaries here aggregate daily
cross sections, never pooled asset/day observations. All dates are naive session
labels, not information-availability timestamps.
"""
from __future__ import annotations
from dataclasses import dataclass
from decimal import Decimal
import math
from numbers import Real
from typing import Any
import numpy as np
import pandas as pd
from scipy import stats
from quant_engine.factor_library import ic_summary, spearman_ic
FACTOR_DIAGNOSTICS_VERSION = "0.1.0"
# IC is bounded to [-1, 1]. Below this absolute float64 resolution, dispersion
# cannot reliably distinguish equivalent cross sections from rounding noise.
MINIMUM_IC_STD_FOR_RATIOS = 32 * np.finfo(np.float64).eps
def _qualification() -> dict[str, Any]:
return {
"contract_version": FACTOR_DIAGNOSTICS_VERSION,
"production_algorithm_version": None,
"source_admission": "not_established",
"historical_availability": "not_established",
"decision_eligible": False,
}
def _integer(value: int, minimum: int, name: str) -> None:
if type(value) is not int or value < minimum:
raise ValueError(f"{name} must be an integer >= {minimum}")
def _label(value: str) -> bool:
return (isinstance(value, str) and 0 < len(value) <= 128
and value == value.strip() and all(char.isprintable() for char in value))
def _dates(index: pd.Index) -> None:
if (not isinstance(index, pd.DatetimeIndex) or index.tz is not None
or index.hasnans or not index.equals(index.normalize())):
raise ValueError("Expected naive midnight session dates")
def _keys(index: pd.Index) -> None:
if (not isinstance(index, pd.MultiIndex) or list(index.names) != ["date", "asset"]
or not index.is_unique):
raise ValueError("Expected unique (date, asset) observation keys")
_dates(index.get_level_values("date"))
if any(not _label(asset) for asset in index.get_level_values("asset")):
raise ValueError("Asset identifiers must be nonempty strings")
def _numbers(series: pd.Series) -> pd.Series:
numbers = []
for value in series:
if value is None or value is pd.NA:
numbers.append(math.nan)
continue
if isinstance(value, (bool, np.bool_)) or not isinstance(value, (Real, Decimal)):
raise ValueError("Values must be real numbers or missing, without coercion")
try:
number = float(value)
except (OverflowError, ValueError) as exc:
raise ValueError("Value cannot be represented as a finite number") from exc
if math.isinf(number):
raise ValueError("Infinite observations are not permitted")
numbers.append(number)
return pd.Series(numbers, index=series.index, name=series.name, dtype=float)
def _panel(frame: pd.DataFrame) -> pd.DataFrame:
if not isinstance(frame, pd.DataFrame):
raise ValueError("Expected a factor DataFrame")
_keys(frame.index)
if not frame.columns.is_unique or any(not _label(name) for name in frame.columns):
raise ValueError("Factor identifiers must be unique nonempty strings")
result = pd.DataFrame(index=frame.index)
for name in frame.columns:
result[name] = _numbers(frame[name])
return result.sort_index()
@dataclass(frozen=True)
class _Correlation:
value: float | None
n_pairs: int
status: str
def _center_scale(series: pd.Series) -> pd.Series:
# Translate first to preserve distinguishable low bits beside a large offset.
# Opposite extreme endpoints may overflow subtraction; only then scale first.
with np.errstate(all="ignore"):
shifted = series - series.iloc[0]
if not bool(np.isfinite(shifted).all()):
bounded = series / series.abs().max()
shifted = bounded - bounded.iloc[0]
return shifted / shifted.abs().max()
def _correlation(left: pd.Series, right: pd.Series, method: str, minimum: int) -> _Correlation:
# Callers already bind the observation domain. Concat still aligns complete
# keys rather than independently deleting missing left and right values.
paired = pd.concat([left.rename("left"), right.rename("right")], axis=1).dropna()
count = len(paired)
if count < minimum:
return _Correlation(None, count, "no_pairs" if count == 0 else "insufficient_pairs")
constant_left = bool((paired["left"] == paired["left"].iloc[0]).all())
constant_right = bool((paired["right"] == paired["right"].iloc[0]).all())
if constant_left or constant_right:
reason = ("constant_both" if constant_left and constant_right else
"constant_left" if constant_left else "constant_right")
return _Correlation(None, count, reason)
# Pearson needs bounded magnitudes to avoid covariance overflow. Rank the
# original observations: scaling could underflow distinct tiny values into
# artificial ties next to a very large outlier.
with np.errstate(all="ignore"):
if method == "pearson":
scaled_left = _center_scale(paired["left"])
scaled_right = _center_scale(paired["right"])
value = float(ic_summary(scaled_left, scaled_right, periods=(1,),
method="pearson").loc[1, "ic_mean"])
else:
value = float(spearman_ic(paired["left"], paired["right"]))
if not math.isfinite(value) or abs(value) > 1 + 1e-12:
return _Correlation(None, count, "numerical_failure")
return _Correlation(max(-1.0, min(1.0, value)), count, "ok")
def _summary(values: list[float | None], minimum: int) -> dict[str, Any]:
available = [value for value in values if value is not None]
count = len(available)
result: dict[str, Any] = {
"valid_days": count, "missing_days": len(values) - count,
"mean": None, "std": None, "ir": None, "t": None, "p": None,
"status": "no_valid_days" if count == 0 else "insufficient_days",
}
if count == 0:
return result
constant = all(value == available[0] for value in available)
mean = available[0] if constant else math.fsum(available) / count
result["mean"] = mean
if count < minimum:
return result
std = 0.0 if constant else float(np.std(available, ddof=1))
result["std"] = std
if constant:
result["status"] = "constant_values"
return result
if not math.isfinite(std) or std <= 0:
result["std"] = None
result["status"] = "numerical_failure"
return result
if std <= MINIMUM_IC_STD_FOR_RATIOS:
result["status"] = "below_resolution"
return result
ir = mean / std
t_value = ir * math.sqrt(count)
p_value = float(2 * stats.t.sf(abs(t_value), df=count - 1))
if not all(math.isfinite(value) for value in (ir, t_value, p_value)):
result["status"] = "numerical_failure"
return result
result.update(ir=ir, t=t_value, p=p_value, status="ok")
return result
def daily_ic(
factors: pd.DataFrame, returns: pd.Series, *, min_pairs: int = 3, min_days: int = 2,
) -> dict[str, Any]:
"""Daily Pearson/average-tie RankIC and unannualized daily-IC summaries.
Factors define the observation domain. Extra return keys are ignored and
counted; missing return keys remain unavailable. Every factor/date is kept,
including zero-pair dates. IID t/p are descriptive only: serial correlation
and overlapping holding intervals are not corrected or admitted.
"""
_integer(min_pairs, 3, "min_pairs")
_integer(min_days, 2, "min_days")
frame = _panel(factors)
if not isinstance(returns, pd.Series):
raise ValueError("Expected a forward-return Series")
_keys(returns.index)
clean_returns = _numbers(returns)
aligned = clean_returns.reindex(frame.index)
rows: list[dict[str, Any]] = []
summaries = []
for factor_id in frame.columns:
points = []
for day, left in frame[factor_id].groupby(level="date", sort=True):
right = aligned.reindex(left.index)
pearson = _correlation(left, right, "pearson", min_pairs)
rank = _correlation(left, right, "spearman", min_pairs)
point = {
"date": day.strftime("%Y-%m-%d"), "factor_id": factor_id,
"ic": pearson.value, "rank_ic": rank.value, "n_pairs": pearson.n_pairs,
"n_observations": len(left), "n_factor": int(left.notna().sum()),
"n_return": int(right.notna().sum()),
"ic_status": pearson.status, "rank_ic_status": rank.status,
}
rows.append(point)
points.append(point)
summary: dict[str, Any] = {"factor_id": factor_id, "n_days": len(points)}
for field in ("ic", "rank_ic"):
summary.update({f"{field}_{key}": value for key, value in
_summary([point[field] for point in points], min_days).items()})
summaries.append(summary)
return {
**_qualification(), "rows": rows, "summary": summaries,
"method": {"aggregation": "equal_weight_daily_cross_sections", "min_pairs": min_pairs,
"min_days": min_days, "standard_deviation_ddof": 1, "ir_annualized": False,
"minimum_ic_std_for_ratios": float(MINIMUM_IC_STD_FOR_RATIOS),
"ic_std_resolution_policy": "absolute_32_float64_eps",
"rank_ties": "average", "t_method": "naive_iid_unadjusted",
"pearson_normalization": "translate_then_scale_with_overflow_fallback",
"p_method": "two_sided_student_t", "t_degrees_of_freedom": "valid_days - 1",
"serial_correlation_adjusted": False, "holding_overlap_adjusted": False,
"observation_domain": "factor_keys", "missing_policy": "pairwise_complete_keys",
"extra_return_keys": len(returns.index.difference(frame.index))},
}
def correlation_matrix(
factors: pd.DataFrame, *, method: str = "pearson", min_pairs: int = 3,
) -> dict[str, Any]:
"""Pool keyed asset/session pairs, not an average of daily correlations.
Larger cross sections contribute more pairs. Pairwise deletion may produce
a non-PSD matrix; this output is not a covariance/risk-matrix contract.
"""
_integer(min_pairs, 3, "min_pairs")
if method not in ("pearson", "spearman"):
raise ValueError("method must be pearson or spearman")
frame = _panel(factors)
names = list(frame.columns)
estimates = [[_correlation(frame[left], frame[right], method, min_pairs)
for right in names] for left in names]
return {
**_qualification(), "factors": names, "n_observations": len(frame),
"matrix": [[estimate.value for estimate in row] for row in estimates],
"n_pairs": [[estimate.n_pairs for estimate in row] for row in estimates],
"status": [[estimate.status for estimate in row] for row in estimates],
"method": {"correlation": method, "aggregation": "pooled_asset_session_pairwise",
"pearson_normalization": "translate_then_scale_with_overflow_fallback",
"min_pairs": min_pairs, "missing_policy": "pairwise_complete_keys",
"rank_ties": "average", "positive_semidefinite_guaranteed": False},
}
@dataclass(frozen=True)
class ForwardReturns:
"""Owned output frames; frozen attributes do not make pandas objects immutable."""
returns: pd.Series
intervals: pd.DataFrame
metadata: dict[str, Any]
def forward_returns(
prices: pd.Series, *, sessions: pd.DatetimeIndex, entry_lag_sessions: int,
holding_sessions: int, price_field: str, price_basis: str,
) -> ForwardReturns:
"""Label P[t+lag+holding]/P[t+lag]-1 on an explicit session calendar.
Requires comparable endpoint prices, not every intermediate price. Missing
keys/prices remain missing on the supplied calendar. lag=0 is a same-session
price basis, not a claim that a closing signal can execute at that close.
"""
_integer(entry_lag_sessions, 0, "entry_lag_sessions")
_integer(holding_sessions, 1, "holding_sessions")
if not _label(price_field) or not _label(price_basis):
raise ValueError("Explicit price field and comparable-price basis are required")
_dates(sessions)
if not sessions.is_unique or not sessions.is_monotonic_increasing:
raise ValueError("Calendar sessions must be unique and increasing")
if not isinstance(prices, pd.Series):
raise ValueError("Expected a keyed price Series")
_keys(prices.index)
values = _numbers(prices)
if len(prices.index.get_level_values("date").difference(sessions)):
raise ValueError("Price observation outside the supplied calendar")
if bool((values.dropna() <= 0).any()):
raise ValueError("Endpoint prices must be positive when present")
assets = sorted(prices.index.get_level_values("asset").unique())
index = pd.MultiIndex.from_product([sessions, assets], names=["date", "asset"])
grid = values.reindex(index)
output, intervals = [], []
for position, signal_date in enumerate(sessions):
entry_position = position + entry_lag_sessions
exit_position = entry_position + holding_sessions
entry_date = sessions[entry_position] if entry_position < len(sessions) else None
exit_date = sessions[exit_position] if exit_position < len(sessions) else None
for asset in assets:
value, status = math.nan, "insufficient_calendar"
if entry_date is not None and exit_date is not None:
entry, exit_price = grid.loc[(entry_date, asset)], grid.loc[(exit_date, asset)]
status = "missing_price"
if pd.notna(entry) and pd.notna(exit_price):
with np.errstate(over="ignore", invalid="ignore"):
value = float(exit_price / entry - 1)
status = "ok" if math.isfinite(value) else "numerical_failure"
if status != "ok":
value = math.nan
output.append(value)
intervals.append({"signal_date": signal_date.strftime("%Y-%m-%d"),
"entry_date": entry_date.strftime("%Y-%m-%d") if entry_date is not None else None,
"exit_date": exit_date.strftime("%Y-%m-%d") if exit_date is not None else None,
"status": status})
metadata = {**_qualification(), "formula": "exit_price / entry_price - 1", "unit": "ratio",
"entry_lag_sessions": entry_lag_sessions, "holding_sessions": holding_sessions,
"price_field": price_field, "price_basis": price_basis,
"price_coverage": "endpoints_only", "calendar": "explicit_caller_sessions",
"missing_policy": "no_fill_no_session_skipping", "execution_eligibility": "not_established"}
return ForwardReturns(pd.Series(output, index=index, dtype=float, name="forward_return"),
pd.DataFrame(intervals, index=index,
columns=["signal_date", "entry_date", "exit_date", "status"]), metadata)
+318
View File
@@ -0,0 +1,318 @@
"""Deterministic diagnostics contracts with independently checkable samples."""
from __future__ import annotations
import importlib
from math import sqrt
import numpy as np
import pandas as pd
import pytest
def api():
return importlib.import_module("quant_engine.factor_diagnostics")
def panel(values, *, days=1, assets=None):
assets = assets or ["A", "B", "C"]
index = pd.MultiIndex.from_product(
[pd.date_range("2026-09-21", periods=days), assets], names=["date", "asset"]
)
return pd.DataFrame(values, index=index)
def test_pairwise_matrix_keeps_complete_security_date_pairs():
frame = panel({"pe": [None, 0, 1, 2, 3, 4, 5, 6],
"pb": [1000, None, 1, 2, 3, 4, 5, 6]}, assets=list("ABCDEFGH"))
original = frame.copy()
result = api().correlation_matrix(frame)
assert result["matrix"][0][1] == pytest.approx(1)
assert result["n_pairs"] == [[7, 6], [6, 7]]
assert result["status"][0][1] == "ok"
assert result["method"]["aggregation"] == "pooled_asset_session_pairwise"
assert result["method"]["positive_semidefinite_guaranteed"] is False
pd.testing.assert_frame_equal(frame, original)
def test_matrix_empty_constant_and_short_diagonals_are_undefined():
frame = panel({"empty": [None] * 3, "constant": [0.05] * 3, "short": [1, 2, None]})
result = api().correlation_matrix(frame)
assert result["matrix"] == [[None] * 3 for _ in range(3)]
assert result["n_pairs"][0][0] == 0
assert result["status"][0][0] == "no_pairs"
assert result["status"][1][1] == "constant_both"
assert result["status"][2][2] == "insufficient_pairs"
def test_rank_correlation_uses_average_ties_after_pairing():
frame = panel({"f": [1, 1, 2, None], "r": [1, 2, 3, 999]}, assets=list("ABCD"))
result = api().correlation_matrix(frame, method="spearman")
assert result["matrix"][0][1] == pytest.approx(sqrt(3) / 2)
assert result["n_pairs"][0][1] == 3
def test_daily_ic_is_not_one_pooled_correlation():
frame = panel({"f": [1, 2, 3, 101, 102, 103]}, days=2)
returns = pd.Series([3, 2, 1, 101, 102, 103], index=frame.index)
result = api().daily_ic(frame, returns)
assert [r["ic"] for r in result["rows"]] == pytest.approx([-1, 1])
assert result["summary"][0]["ic_mean"] == pytest.approx(0)
assert result["summary"][0]["ic_std"] == pytest.approx(sqrt(2))
assert result["summary"][0]["ic_ir"] == 0
assert result["summary"][0]["ic_t"] == 0
assert result["method"]["ir_annualized"] is False
assert result["method"]["t_method"] == "naive_iid_unadjusted"
assert result["method"]["serial_correlation_adjusted"] is False
def test_daily_summary_matches_three_known_days_and_keeps_zero():
frame = panel({"f": [1, 2, 3] * 3}, days=3)
returns = pd.Series([3, 2, 1, 1, 0, 1, 1, 2, 3], index=frame.index)
result = api().daily_ic(frame, returns)
assert [r["ic"] for r in result["rows"]] == pytest.approx([-1, 0, 1])
assert [r["rank_ic"] for r in result["rows"]] == pytest.approx([-1, 0, 1])
summary = result["summary"][0]
assert (summary["n_days"], summary["ic_valid_days"], summary["rank_ic_valid_days"]) == (3, 3, 3)
assert summary["ic_std"] == pytest.approx(1)
assert summary["ic_ir"] == pytest.approx(0, abs=1e-15)
assert summary["ic_t"] == pytest.approx(0, abs=1e-15)
assert summary["rank_ic_ir"] == pytest.approx(0, abs=1e-15)
assert summary["rank_ic_t"] == pytest.approx(0, abs=1e-15)
def test_undefined_days_and_factor_observation_domain_survive_alignment():
frame = panel({"f": [1, 2, 3, None, None, None, 1, 2, 3]}, days=3)
returns = pd.Series([1, 2, 3], index=frame.index[:3])
extra = pd.Series([100], index=pd.MultiIndex.from_tuples(
[(pd.Timestamp("2026-09-25"), "D")], names=["date", "asset"]))
result = api().daily_ic(frame, pd.concat([returns, extra]))
assert len(result["rows"]) == 3
assert [r["n_pairs"] for r in result["rows"]] == [3, 0, 0]
assert [r["n_observations"] for r in result["rows"]] == [3, 3, 3]
assert [r["ic_status"] for r in result["rows"]] == ["ok", "no_pairs", "no_pairs"]
summary = result["summary"][0]
assert summary["ic_valid_days"] == 1
assert summary["ic_missing_days"] == 2
assert summary["ic_mean"] == pytest.approx(1)
assert summary["ic_std"] is summary["ic_ir"] is summary["ic_t"] is None
assert summary["ic_status"] == "insufficient_days"
assert result["method"]["extra_return_keys"] == 1
def test_daily_summary_of_constant_ic_has_no_ratio_statistics():
frame = panel({"f": [1, 2, 3] * 3}, days=3)
returns = pd.Series([1, 1, 2] * 3, index=frame.index)
summary = api().daily_ic(frame, returns)["summary"][0]
assert summary["ic_mean"] == pytest.approx(sqrt(3) / 2)
assert summary["ic_std"] == 0
assert summary["ic_ir"] is summary["ic_t"] is None
assert summary["ic_status"] == "constant_values"
assert summary["rank_ic_ir"] is summary["rank_ic_t"] is None
def test_permutation_and_input_copies_do_not_change_diagnostics():
frame = panel({"f": [1, 2, 3, 4, 5, 6]}, days=2)
returns = pd.Series([1, 3, 2, 4, 6, 5], index=frame.index)
original_frame, original_returns = frame.copy(), returns.copy()
baseline = api().daily_ic(frame, returns)
assert api().daily_ic(frame.iloc[::-1], returns.iloc[[2, 0, 4, 1, 5, 3]]) == baseline
pd.testing.assert_frame_equal(frame, original_frame)
pd.testing.assert_series_equal(returns, original_returns)
def test_large_finite_inputs_reuse_scaled_correlation_without_overflow():
frame = panel({"a": [-1e308, 0, 1e308], "b": [-1e307, 0, 1e307]})
assert api().correlation_matrix(frame)["matrix"][0][1] == pytest.approx(1)
@pytest.mark.parametrize("value", [True, "1.2", np.inf, -np.inf, object()])
def test_diagnostics_reject_invalid_numbers(value):
frame = panel({"f": [value, 2, 3]})
with pytest.raises(ValueError, match=r"Values|Infinite"):
api().correlation_matrix(frame)
@pytest.mark.parametrize("kind", ["duplicate_keys", "bad_names", "bad_date", "timezone", "empty_asset", "duplicate_factors"])
def test_diagnostics_reject_ambiguous_keys(kind):
frame = panel({"f": [1, 2, 3]})
if kind == "duplicate_keys":
frame.index = pd.MultiIndex.from_tuples([frame.index[0]] * 3, names=["date", "asset"])
elif kind == "bad_names":
frame.index.names = ["day", "asset"]
elif kind == "bad_date":
frame.index = pd.MultiIndex.from_product([["2026-09-21"], list("ABC")], names=["date", "asset"])
elif kind == "timezone":
frame.index = pd.MultiIndex.from_product([pd.date_range("2026-09-21", periods=1, tz="UTC"), list("ABC")], names=["date", "asset"])
elif kind == "empty_asset":
frame.index = pd.MultiIndex.from_product([pd.date_range("2026-09-21", periods=1), ["", "B", "C"]], names=["date", "asset"])
else:
frame = pd.concat([frame, frame], axis=1)
with pytest.raises(ValueError, match=r"keys|dates|identifiers"):
api().correlation_matrix(frame)
@pytest.mark.parametrize("minimum", [True, 0, 2, 3.1, "3"])
def test_minimum_pairs_is_explicit_and_at_least_three(minimum):
with pytest.raises(ValueError, match="min_pairs"):
api().correlation_matrix(panel({"f": [1, 2, 3]}), min_pairs=minimum)
def test_all_missing_and_empty_inputs_remain_observed_not_identity_or_zero():
frame = panel({"f": [None] * 3})
result = api().daily_ic(frame, pd.Series([None] * 3, index=frame.index))
assert result["summary"][0]["ic_mean"] is None
assert result["summary"][0]["ic_status"] == "no_valid_days"
empty = frame.iloc[:0]
assert api().correlation_matrix(empty)["matrix"] == [[None]]
assert api().daily_ic(empty, pd.Series([], index=empty.index, dtype=float))["rows"] == []
def prices(values=(100, 110, 121, 133.1)):
sessions = pd.DatetimeIndex(["2026-09-18", "2026-09-21", "2026-09-22", "2026-09-23"])
index = pd.MultiIndex.from_product([sessions, ["A"]], names=["date", "asset"])
return pd.Series(values, index=index), sessions
def forward(series, sessions, **changes):
options = {"entry_lag_sessions": 0, "holding_sessions": 2,
"price_field": "close", "price_basis": "synthetic_comparable"}
options.update(changes)
return api().forward_returns(series, sessions=sessions, **options)
def test_forward_returns_use_explicit_calendar_endpoints_and_compound_price_ratio():
series, sessions = prices()
result = forward(series, sessions)
assert result.returns.iloc[0] == pytest.approx(0.21)
assert result.intervals.iloc[0]["entry_date"] == "2026-09-18"
assert result.intervals.iloc[0]["exit_date"] == "2026-09-22"
assert result.intervals.iloc[-1]["status"] == "insufficient_calendar"
assert np.isnan(result.returns.iloc[-1])
assert result.metadata["formula"] == "exit_price / entry_price - 1"
assert result.metadata["entry_lag_sessions"] == 0
assert result.metadata["holding_sessions"] == 2
assert result.metadata["execution_eligibility"] == "not_established"
assert result.metadata["historical_availability"] == "not_established"
def test_forward_entry_lag_changes_both_endpoints():
series, sessions = prices([10, 20, 30, 80])
result = forward(series, sessions, entry_lag_sessions=1)
assert result.returns.iloc[0] == pytest.approx(3)
assert result.intervals.iloc[0]["entry_date"] == "2026-09-21"
assert result.intervals.iloc[0]["exit_date"] == "2026-09-23"
def test_forward_does_not_skip_missing_prices_or_derive_calendar_from_rows():
series, sessions = prices([100, None, 121, 133.1])
result = forward(series.drop(index=(sessions[1], "A")), sessions, holding_sessions=1)
assert len(result.returns) == 4
assert np.isnan(result.returns.iloc[0])
assert np.isnan(result.returns.iloc[1])
assert result.intervals.iloc[0]["status"] == "missing_price"
assert result.intervals.iloc[0]["exit_date"] == "2026-09-21"
assert result.returns.iloc[2] == pytest.approx(0.1)
@pytest.mark.parametrize("options", [{"holding_sessions": 0}, {"holding_sessions": True},
{"holding_sessions": 1.5}, {"entry_lag_sessions": -1},
{"entry_lag_sessions": False}, {"price_field": ""},
{"price_basis": ""}])
def test_forward_rejects_ambiguous_interval_parameters(options):
series, sessions = prices()
with pytest.raises(ValueError, match=r"sessions|price field"):
forward(series, sessions, **options)
@pytest.mark.parametrize("kind", ["duplicate", "descending", "timezone", "intraday", "outside"])
def test_forward_rejects_invalid_calendar(kind):
series, sessions = prices()
if kind == "duplicate":
sessions = sessions.insert(1, sessions[0])
elif kind == "descending":
sessions = sessions[::-1]
elif kind == "timezone":
sessions = sessions.tz_localize("UTC")
elif kind == "intraday":
sessions = sessions + pd.Timedelta(hours=1)
else:
sessions = sessions[:-1]
with pytest.raises(ValueError, match=r"sessions|dates|calendar"):
forward(series, sessions)
@pytest.mark.parametrize("value", [0, -1, True, "2", np.inf])
def test_forward_rejects_nonpositive_or_nonfinite_prices(value):
series, sessions = prices([value, 110, 121, 133.1])
with pytest.raises(ValueError, match=r"positive|Values|Infinite"):
forward(series, sessions)
def test_result_version_and_parameters_do_not_grant_source_or_decision_admission():
frame = panel({"f": [1, 2, 3]})
result = api().daily_ic(frame, pd.Series([1, 2, 3], index=frame.index))
assert result["contract_version"] == api().FACTOR_DIAGNOSTICS_VERSION
assert result["method"]["min_pairs"] == 3
assert result["method"]["min_days"] == 2
assert result["production_algorithm_version"] is None
assert result["source_admission"] == result["historical_availability"] == "not_established"
assert result["decision_eligible"] is False
def test_nonzero_daily_summary_uses_three_days_not_nine_asset_rows():
frame = panel({"f": [1, 2, 3] * 3}, days=3)
returns = pd.Series([.01, .02, .03, .02, .03, .01, .01, -.02, .01], index=frame.index)
result = api().daily_ic(frame, returns)
assert [point["ic"] for point in result["rows"]] == pytest.approx([1, -.5, 0], abs=1e-15)
summary = result["summary"][0]
for prefix in ("ic", "rank_ic"):
assert summary[f"{prefix}_valid_days"] == 3
assert summary[f"{prefix}_mean"] == pytest.approx(1 / 6)
assert summary[f"{prefix}_std"] == pytest.approx(sqrt(7 / 12))
assert summary[f"{prefix}_ir"] == pytest.approx(.2182178902359924)
assert summary[f"{prefix}_t"] == pytest.approx(.3779644730092272)
assert summary[f"{prefix}_p"] == pytest.approx(.7418011102528389)
def test_each_side_has_three_values_but_only_two_common_pairs():
frame = panel({"left": [1, 2, 3, None], "right": [None, 2, 3, 4]}, assets=list("ABCD"))
result = api().correlation_matrix(frame)
assert result["n_pairs"][0][1] == 2
assert result["matrix"][0][1] is None
assert result["status"][0][1] == "insufficient_pairs"
def test_rank_preserves_distinct_tiny_values_beside_a_large_outlier():
frame = panel({"left": [1e-308, 2e-308, 1e308], "right": [1, 2, 3]})
assert api().correlation_matrix(frame, method="spearman")["matrix"][0][1] == pytest.approx(1)
def test_daily_rank_ic_preserves_tiny_distinct_observations():
frame = panel({"f": [1e-200, 2e-200, 3e-200, 1e308]}, assets=list("ABCD"))
returns = pd.Series([4, 3, 2, 1], index=frame.index)
result = api().daily_ic(frame, returns)
assert result["rows"][0]["rank_ic"] == pytest.approx(-1)
assert result["rows"][0]["rank_ic_status"] == "ok"
def test_affine_equivalent_daily_ic_does_not_turn_roundoff_into_extreme_significance():
frame = panel({"f": [1, 2, 3, 10, 20, 30, 101, 102, 103]}, days=3)
returns = pd.Series([1, 1, 2, 101, 101, 102, 100001, 100001, 100002], index=frame.index)
result = api().daily_ic(frame, returns)
assert [point["ic"] for point in result["rows"]] == pytest.approx([sqrt(3) / 2] * 3)
summary = result["summary"][0]
assert summary["ic_mean"] == pytest.approx(sqrt(3) / 2)
assert summary["ic_valid_days"] == 3
assert summary["ic_ir"] is summary["ic_t"] is summary["ic_p"] is None
assert summary["ic_std"] <= result["method"]["minimum_ic_std_for_ratios"]
assert summary["ic_status"] in {"constant_values", "below_resolution"}
@pytest.mark.parametrize("offset,step", [(1e16, 2.0), (1e12, np.spacing(1e12))])
def test_pearson_preserves_representable_differences_beside_a_large_offset(offset, step):
frame = panel({"f": offset + np.arange(5) * step}, assets=list("ABCDE"))
returns = pd.Series([2, 5, 1, 4, 3], index=frame.index)
matrix = api().correlation_matrix(frame.assign(other=returns))
assert matrix["matrix"][0][1] == pytest.approx(.1, abs=1e-14)
daily = api().daily_ic(frame, returns)
assert daily["rows"][0]["ic"] == pytest.approx(.1, abs=1e-14)